YOLOv8 Explained: Real-Time Object Detection in 2026

YOLOv8 Explained: Real-Time Object Detection in 2026

Picture a camera above a supermarket checkout lane, producing about 30 pictures every second. Software has to find every person and trolley in each picture, draw a box around each, and pass the result along before the next picture arrives. If it takes too long, frames pile up and the numbers stop meaning anything.

That job is called real-time object detection, and for the past few years a large share of the teams doing it have reached for the same tool first: YOLOv8, released by the company Ultralytics in January 2023.

In 2026, YOLOv8 is no longer the newest model in its family. Ultralytics shipped YOLO11 in 2024 and YOLO26 in January 2026, and both beat it in benchmark tests. Even so, YOLOv8 still runs inside plenty of shipped products, and most online courses were written for it. This guide explains how it works, where it goes wrong in real deployments, and how to decide whether it belongs in your product.

Takeaways before you read on

1.      YOLOv8 finds objects in an image and returns a box, a label and a confidence score for each one, fast enough to keep up with live video on modest hardware.

2.      It comes in five sizes, from nano (3.2 million parameters) to extra-large (68.2 million). Start small and move up only if needed.

3.      YOLO26 is faster on ordinary processors and more accurate, but YOLOv8 is stable and heavily documented, and code written for one mostly works with the other.

4.      The model is rarely the weakest part of a project. Missing training data and awkward camera angles cause most failures.

5.      The Ultralytics code is licensed under AGPL-3.0. Closed-source commercial products usually need an enterprise license.

What YOLOv8 actually does

Hand YOLOv8 a photo and it hands back a list. Each entry answers three questions: what is this, where is it, and how sure are you? The "where" is a rectangle described by four numbers. The "what" is a label such as person, forklift or cardboard box. The "how sure" is a confidence score between 0 and 1. One line of output might read: person, 0.91, at these coordinates.

That is what YOLOv8 object detection means in practice. The model has no idea that the person looks impatient. It spots things it was trained to spot and reports where they are. Anything smart a product does afterwards, like counting or sending an alert, is ordinary code written on top of that list.

The name comes from a 2015 research paper by Joseph Redmon and colleagues called "You Only Look Once." Detectors of that era first proposed thousands of regions that might contain an object, then classified each region separately. YOLO folded both steps into a single pass through one neural network, which is why the family became known for speed.

A computer vision model is a program that has learned to recognize patterns in pixels by studying a large collection of labeled images. YOLOv8 ships already trained on COCO, a public dataset with 80 everyday categories such as people, cars, bicycles, dogs, cups, chairs and laptops. You can use it as it comes, or retrain it to recognize things specific to your business, like a crack in a weld.

The version numbers trip people up. Ultralytics built YOLOv8, YOLO11 and YOLO26, while YOLOv9 and YOLOv10 came from separate academic groups. A higher number does not automatically mean a better fit for your use.

How the model looks at an image

You don't need the internals to use YOLOv8, but a rough mental picture helps when results look wrong. The network has three sections.

The backbone comes first. It pulls out visual features at several scales, from edges and textures to shapes and object parts. This stage notices "something round and shiny is here" without deciding what it is.

The neck mixes those scales together. Small objects show up best in the fine, detailed layers, while large objects make more sense in the coarse ones. The neck passes information between them so a distant pedestrian and a nearby bus can both be found.

The head makes the final call. For each cell of a grid laid over the image, it predicts whether an object is centered there, its class, and how far its box stretches to each edge.

Three design choices separate YOLOv8 from the YOLO versions before it. The first is that it is anchor-free. Older versions started from a fixed menu of box shapes, called anchors, and nudged them to fit each object. Objects with odd proportions, like long pipes, fit those presets badly. YOLOv8 predicts the center point and distance to each edge directly, so no anchor tuning is needed.

The second is a split head. One branch answers "what is it?" and another answers "where exactly is it?" The two questions need different kinds of features, so keeping them apart helps both. The third is a backbone block called C2f, which sends information along parallel paths before merging it. You won't configure it, but it's part of why a small YOLOv8 network stays accurate.

One more step happens after the network finishes, and it matters later in this article. The model makes many overlapping guesses for the same object. A cleanup rule called non-maximum suppression, or NMS, keeps the most confident box and deletes other boxes that overlap it by more than a set amount. Overlap is measured with IoU (intersection over union), the area two boxes share divided by the total area they cover. NMS is a hand-written rule outside the neural network. It slows down in crowded scenes and can delete real objects standing close together. The newest Ultralytics model, YOLO26, was designed to get rid of it.

Five sizes and five jobs

YOLOv8 comes in five sizes: nano (n), small (s), medium (m), large (l) and extra-large (x). Bigger models are more accurate and slower. The official benchmark figures show how steep that trade-off is.

Model

Parameters

Accuracy (COCO mAP 50-95)

Time per image, CPU

Time per image, A100 GPU

YOLOv8n

3.2 million

37.3

80.4 ms

0.99 ms

YOLOv8s

11.2 million

44.9

128.4 ms

1.20 ms

YOLOv8m

25.9 million

50.2

234.7 ms

1.83 ms

YOLOv8l

43.7 million

52.9

375.2 ms

2.39 ms

YOLOv8x

68.2 million

53.9

479.1 ms

3.53 ms

Source: Ultralytics YOLOv8 documentation. Images at 640 pixels. CPU times use ONNX Runtime; GPU times use TensorRT. Measured on an Amazon EC2 P4d instance.

Two terms in that table need a quick explanation. Parameters are the adjustable numbers inside the network set during training; more of them means more memory and computing. mAP, short for mean average precision, is a score out of 100 that measures how closely the predicted boxes match the real ones across a range of strictness levels. A jump from 37 to 54 sounds modest, but in practice it separates a model that misses many small or partly hidden objects from one that catches most of them.

Watch the CPU column if you're planning a budget. The nano model handles roughly 12 frames a second on a server processor; the extra-large model manages about two. On a data-center GPU, every size finishes in under four milliseconds. So whether YOLOv8 object detection counts as "fast" depends almost entirely on which size you pick and what hardware runs it. A laptop and a cloud GPU will give very different numbers from the same model file.

Detection is only one of the jobs the package handles. The same library offers:

•         Detection, which draws a labeled box around each object.

•         Instance segmentation, which traces each object's exact outline pixel by pixel.

•         Classification, which gives one label to the whole image, for example "defective" or "fine."

•         Pose estimation, which finds 17 body points on each person, such as shoulders, elbows and knees.

•         Oriented bounding boxes (OBB), which are rotated boxes for objects seen at an angle, like ships in satellite photos or papers lying on a desk.

A built-in tracking mode, using trackers called ByteTrack and BoT-SORT, gives each object an ID that stays with it from frame to frame. Tracking turns "there are four people in this frame" into "this particular person has been in the queue for three minutes."

Pro tip

Start with the nano or small model, even if you expect to need more accuracy. Train it on your own images, look at where it fails, and fix the data first. Teams that jump straight to the extra-large model often learn later that the real problem was their labels.

Why money keeps moving into vision

Research firms agree computer vision spending is growing fast. They disagree widely about its current size.

Research firm

2026 market estimate

Forecast

Yearly growth rate

Grand View Research (July 2026)

$28.2 billion

$101.5 billion by 2033

20.1%

Mordor Intelligence (2026)

$32.88 billion

$68.38 billion by 2031

15.77%

The Business Research Company (2026)

$20.52 billion

$37.1 billion by 2030

15.9%

That is a gap of more than $12 billion for the same year. The difference comes from definitions: some firms count cameras, sensors and chips heavily, while others weight software more. Grand View Research found that hardware made up 71.5% of revenue in 2025, with software as the fastest-growing piece. Mordor Intelligence expects automotive uses to grow fastest, at 18.23% a year through 2031, as cars carry more cameras.

For a business reader, a computer vision model like YOLOv8 sits in the software slice, which is smaller than hardware but growing faster. The shift toward running models on cameras and small devices is exactly the work YOLO models were built for.

YOLOv8 next to the newer models

By 2026, YOLOv8 has been overtaken twice by its own maker.

YOLO11, released in September 2024, swapped some building blocks for more efficient ones and added an attention module, which helps the network focus on the most informative parts of the image.

YOLO26, launched on January 14, 2026, made deeper changes. It removed NMS, so the network outputs final boxes directly. It also dropped a component called Distribution Focal Loss (DFL), which YOLOv8 uses to sharpen box edges but which makes exporting the model to phones and small chips harder. It also added training techniques aimed at small objects. In the company's own tests, the YOLO26 nano model scores 40.9 on COCO and takes 38.9 milliseconds per image on a CPU, against 37.3 and 80.4 milliseconds for YOLOv8 nano.

Outside Ultralytics, YOLOv10 from Tsinghua University (May 2024) was the first numbered YOLO to skip NMS. Ultralytics says a YOLO27 is in final development, with no launch date or public model files yet.

 

YOLOv8

YOLO11

YOLO26

Released

January 2023

September 2024

January 2026

Uses the NMS cleanup step

Yes

Yes

No

Nano model accuracy (COCO mAP)

37.3

39.5

40.9

Nano model CPU time per image

80.4 ms

56.1 ms

38.9 ms

Nano model parameters

3.2 million

2.6 million

2.4 million

Tasks supported

Detect, segment, classify, pose, OBB

Same five

Same five, plus semantic segmentation and depth estimation

Tutorials and community answers

Largest body of material

Large and growing

Newest, still catching up

Best fit

Existing systems, learning, stable pipelines

Mature production upgrade

New builds on CPUs and edge devices

Accuracy and speed figures from Ultralytics documentation, 640-pixel images, CPU times with ONNX Runtime.

So why would anyone still choose YOLOv8? A working system is worth more than a benchmark gain. If your YOLOv8 pipeline already meets its targets, migrating costs engineering time for a gain users may never notice. The volume of written help also matters: most forum answers and courses from 2023 through 2025 assume YOLOv8. And some hardware toolchains were tested against YOLOv8's exact output format, which changes in YOLO26 because NMS is gone.

All three share one Python interface, so swapping "yolov8n.pt" for "yolo26n.pt" starts a migration, though it rarely finishes one.

For real-time object detection in a new project that has to run on ordinary processors, YOLO26 is the sensible default in 2026. For existing systems and for learning, YOLOv8 still earns its place.

The problems the demo never shows

A demo video of YOLOv8 boxing cars on a busy street takes an afternoon. A system that works on your street, at night, in the rain, for a year, takes much longer.

Data gaps: the model only knows what it has seen

The pretrained model knows 80 COCO categories. It has never seen your product, your factory floor or your camera angle. Even for classes it knows, like "person," it learned mostly from photos taken at eye level in decent light. A ceiling camera looking straight down turns people into dark circles with shoulders, and accuracy drops.

Training on your own images fixes this, but your dataset covers only the conditions someone remembered to photograph. Ultralytics' own training advice suggests at least 1,500 images and 10,000 labeled objects per class, plus a small share of background images (up to about 10%) that contain no objects at all, so the model learns what "nothing here" looks like. Most first projects start with far fewer, and most failures in YOLOv8 object detection projects trace back to that shortfall.

The gaps that hurt most are the ones nobody noticed during planning:

•         Night shifts, when every training photo was taken in the afternoon.

•         Seasonal change, such as winter coats that hide body shapes or snow covering floor and road markings.

•         New packaging, when a supplier redesigns a box and the model stops recognizing it.

•         Rare events that matter a lot, like a spill or a fallen worker, which appear in only a handful of frames.

The fix is a habit: log low-confidence detections in production, review a sample weekly, label the confusing ones and retrain. A model that is never retrained drifts quietly out of date.

Conflicting signals: when outputs contradict each other

YOLOv8 can be confidently wrong, and it can disagree with itself.

Take a camera in a clothing store. In one frame the model labels an object "person" at 0.62 confidence and, on an overlapping box, "mannequin" at 0.58. (Imagine a custom model trained on both classes.) By default, YOLOv8 runs NMS separately for each class, so both boxes survive, and your counting code now sees two objects where there is one. Turning on class-agnostic NMS (the agnostic_nms=True setting) makes the cleanup compare boxes across classes and keep only the stronger one. That solves this case and creates a new one: two real objects of different types standing close together can now wipe each other out.

Flicker is the other common conflict: a pallet is detected in frame 101, missed in 102 and found in 103, so a "disappeared" alert fires constantly. Trackers help by keeping an object alive for a few frames after it vanishes, and most production systems add a voting rule, such as "the object must appear in 5 of the last 8 frames before we act." That costs a fraction of a second and removes many false alarms.

The confidence threshold defaults to 0.25. Raise it and you get fewer false alarms but more misses; lower it and the opposite happens. No single value is right. A safety system watching for people near moving machinery should accept extra false alarms, while a shopper-counting dashboard can live with a few misses. Set it by what each mistake costs, and if costs differ between classes, filter each class with its own cut-off after prediction.

Real-time decisions: working inside the frame budget

A camera running at 30 frames per second gives you about 33 milliseconds per frame, and the model is only one part of what must fit inside that window.

Stage

What happens

Rough time

Grab and decode

Pull the frame from the camera stream and unpack it

3 to 8 ms

Resize and prepare

Shrink the image to 640 pixels and convert colors

1 to 3 ms

Run the model

The network produces raw predictions

10 to 25 ms

NMS cleanup

Remove duplicate boxes

1 to 10 ms, longer in crowds

Your own logic

Tracking, counting, rules and alerts

1 to 5 ms

Illustrative figures for a small model on an edge GPU at 30 frames per second. Real numbers vary widely by hardware, so measure your own pipeline.

Add up the high end of each row and you've blown the budget. You can run detection on every second or third frame and let the tracker fill the gaps, shrink the input from 640 pixels to 480 or 320 (which hurts small objects), or move to a smaller model. Never let frames queue up. A detector slower than its camera falls further behind every second and soon reports events that are long over. For real-time object detection, dropping a stale frame is nearly always better than processing it late.

Look again at the NMS row. Its cost grows with the number of candidate boxes, so your worst delays arrive exactly when a packed entrance gives you the most people to watch. That unpredictability is one reason Ultralytics removed the step in YOLO26.

Exceptions and edge cases

Small objects are the most common complaint. Shrink a 4K frame to 640 pixels and a distant person may end up a few pixels tall. You can train and run at a larger input size such as 1280 pixels, or cut the image into overlapping tiles and run detection on each tile (the open-source SAHI library is often used for this). Both cost extra computing time.

Crowds cause a different problem. When people stand shoulder to shoulder, their boxes overlap heavily, and NMS may merge two people into one. Lowering the IoU threshold makes the merging worse, while raising it lets more duplicate boxes through. Dense crowd counting often needs a different method, such as density estimation, which predicts a head count for an area without boxing each person.

Pictures of things fool the model regularly. A person on a billboard or a car reflected in a shop window looks real at the pixel level. Add such images to training as background, or mark zones in the frame where detections are ignored.

Motion blur smears fast-moving objects indoors, and raising the camera's shutter speed often helps more than any model change.

Label disagreement is the sneakiest edge case. Is a box with one torn corner "damaged"? If two labelers would answer differently, the model learns the inconsistency. Write a one-page labeling guide with example images before anyone starts.

Under pressure: what changes at scale

Two hundred cameras across forty stores is a different project from one camera on a laptop. The first thing to strain is usually hardware throughput, not accuracy. A single GPU can serve many camera streams if you batch frames, meaning you feed several images through the model at once. Batching raises throughput but adds a small wait while each batch fills, so keep batches small for safety alerts.

Exporting to an optimized format helps a lot. The package can convert YOLOv8 to ONNX, TensorRT (for NVIDIA GPUs), OpenVINO (for Intel chips), CoreML (for Apple devices) and TFLite (for Android phones and small boards). TensorRT can also run the model with smaller number formats (FP16 or INT8) for extra speed. INT8 needs a calibration step and can cost some accuracy, so test it on your own footage.

Bandwidth is the quiet cost. Many large deployments run a small model on or near each camera and send only the results, a few hundred bytes of boxes and labels, instead of megabytes of video. Raw footage then never leaves the building, which also helps with privacy rules.

At this scale, a computer vision model needs the same care as any other production software. Each trained version needs a number, a record of its dataset, a way to roll back, and monitoring of confidence scores. A slow drop in average confidence is often the first sign that a camera has been bumped, a lens is dirty or the scene itself has changed.

Plan for failure too. Cameras go offline and networks drop. Decide ahead of time what the system does when it's blind. Silent failure is the worst outcome, because the dashboard keeps showing zero incidents while nothing is being watched.

Pro tip

Keep a folder of 200 to 500 hard images taken from real production footage: night shots, crowded moments, reflections, blurred frames. Run every new model version against that folder before release. It will tell you more about real performance than the COCO score ever will.

A short YOLOv8 tutorial: from install to your first custom model

This YOLOv8 tutorial assumes you have Python and a terminal. A GPU is optional, though CPU training is slow.

Step 1: Install the package.

pip install ultralytics

Step 2: Run the pretrained model on an image.

from ultralytics import YOLO

 

model = YOLO("yolov8n.pt")

results = model("street.jpg", conf=0.25)

for box in results[0].boxes:

print(model.names[int(box.cls)], float(box.conf))

The model file downloads on first use, and you'll see lines such as "person 0.89."

Step 3: Prepare your own data.

Each image needs a matching text file with one line per object: the class number, then the box's center position, width and height, all written as fractions of the image size. Labeling tools such as CVAT and Roboflow export this format. A small settings file tells the trainer where the images are and what the classes are called:

path: datasets/parcels

train: images/train

val: images/val

names:

  0: parcel

  1: damaged_parcel

Step 4: Train.

model = YOLO("yolov8n.pt")

model.train(data="parcels.yaml", epochs=100, imgsz=640)

Starting from pretrained weights, called transfer learning, means the model already recognizes edges and shapes and only has to learn your categories. By default, training uses mosaic augmentation, which stitches four images into one for extra variety, and turns it off for the final 10 epochs (full passes through the data) so the model finishes on normal-looking images.

Step 5: Check the results.

Training saves charts and scores in a folder called runs. Look past the single mAP number and open the confusion matrix, which shows which classes get mistaken for which. If "damaged_parcel" keeps getting labeled "parcel," you need more damaged examples, and clearer ones.

Step 6: Export for deployment.

model.export(format="onnx")

Replace "onnx" with "engine" for TensorRT, or with "openvino", "coreml" or "tflite", to match your hardware.

Those six steps cover the core of any YOLOv8 tutorial. To try the newer model later, change "yolov8n.pt" to "yolo26n.pt", rerun the same steps, and compare both models on your folder of hard images.

Pro tip

Split your data by time or location, not at random. If frames from the same video clip land in both the training and validation sets, the validation score will look far better than real-world performance, because the model has effectively seen the answers already.

A licensing detail founders often miss

The Ultralytics code and model files are released under the AGPL-3.0 license. Put simply, if you build software with it and offer that software to other people, including as a web service, you may have to publish your own source code under the same license. Many companies building closed commercial products buy an Ultralytics Enterprise License instead. Check this before you ship, and talk to a lawyer if your situation is unclear.

Is YOLOv8 right for your team?

The answer depends on where you sit.

Founders scoping a feature should know the model is the cheap part. Most of the budget goes into collecting and labeling real images, choosing hardware, and the retraining loop after launch. A proof of concept can take days; a dependable product usually takes months.

Developers will find YOLOv8 a gentle way into object detection. For a new production build, prototype with both YOLOv8 and YOLO26, since the code is nearly identical.

Operations and office teams should start from the question. "How many people used meeting room B today?" suits a detector. "Was the meeting useful?" does not. Tasks that come down to spotting, counting or locating visible things are a good match for YOLOv8 object detection. Tasks that need judgment, context or a sense of intent are not.

Content writers should be careful with version claims: "YOLOv8 is the latest YOLO" stopped being true in 2024. A claim that one model is "faster" means little without the size and hardware attached.

Where this leaves you

YOLOv8 turned a hard research problem into something one developer can set up in an afternoon, which is why it spread so widely. In 2026 it is no longer the best performer in its family. YOLO26 beats it on ordinary processors, removes the NMS step that slows down crowded scenes, and keeps the same Python interface, which makes it the natural starting point for new edge projects.

Still, the model choice matters less than most buyers expect. Projects succeed when teams collect data from real conditions, set thresholds based on what mistakes actually cost, plan their frame budget, and keep retraining as the world in front of the camera changes. A team that does those things will get good results from YOLOv8 or YOLO26. A team that skips them will struggle with either.

Ayush Kanodia

Ayush Kanodia

Ayush Kanodia, an esteemed Director at HireFullStackDeveloperIndia, channels his passion into delivering cutting-edge IT services and solutions. Through his leadership, he has driven numerous successful projects, solidifying the company's standing as a pioneering force in the industry.

Build Your Agile Team

We provide you with a top-performing extended team for all your development needs in any technology.

Hourly
$20
It Includes
Duration
Hourly Basis
Communication
Phone, Skype, Slack, Chat, Email
Hiring Period
25 Hours (MIN)
Project Trackers
Daily Reports, Basecamp, Jira, Redmime, etc
Methodology
Agile
Monthly
$2600
It Includes
Duration
160 Hours
Communication
Phone, Skype, Slack, Chat, Email
Hiring Period
1 Month
Project Trackers
Daily Reports, Basecamp, Jira, Redmime, etc
Methodology
Agile
Team
$13200
It Includes
Team Members
1 (PM), 1 (QA), 4 (Developers)
Communication
Phone, Skype, Slack, Chat, Email
Hiring Period
1 Month
Project Trackers
Daily Reports, Basecamp, Jira, Redmime, etc
Methodology
Agile

Frequently Asked Questions

Is YOLOv8 still worth learning in 2026?
Yes. The concepts, the dataset format and the Python interface carry over almost unchanged to YOLO11 and YOLO26, so the time isn't wasted. Many production systems still run YOLOv8, too.
Can YOLOv8 run without a GPU?
Yes, as long as you pick the size carefully. On a server processor, the nano model handles an image in about 80 milliseconds, or roughly 12 frames per second. That works for many real-time object detection jobs that don't need full video frame rates, such as checking a shelf once a second. The larger sizes are generally too slow for live video on a CPU.
How many images do I need to train it on my own objects?
A few hundred labeled images per class can give you a rough working model if you start from pretrained weights. For production quality, Ultralytics suggests around 1,500 images and 10,000 labeled objects per class, covering the lighting, angles and conditions the model will actually face.
What is the main difference between YOLOv8 and YOLO26?
YOLO26 outputs final boxes directly, without the NMS cleanup step, and drops the DFL component that made exporting harder. That gives it up to 43% faster CPU inference in Ultralytics' tests, along with higher accuracy at each size. Your training code barely changes between the two.
Where can I find a good YOLOv8 tutorial beyond this one?
The official Ultralytics documentation has current step-by-step guides for training, prediction, tracking and export. After that, the most useful YOLOv8 tutorial is your own data: label 200 images from your real cameras, train the nano model, and study where it fails.