Technical explainers

How Face Detection and Face Tracking Work Together in Video

Understand face detection, tracking, grouping, and rendering through a simple moving-person example. Diagnose why masks disappear, drift, or switch people.

On this page

Face detection finds faces in an image. Face tracking follows their locations through a sequence. A video anonymizer needs both: finding a face once is not enough to keep a blur attached while the person moves.

Selective anonymization adds another task: associating appearances with the person you chose in the face list.

Follow one person across a room

Imagine a person entering from the left, walking behind a pillar, and returning on the right.

The detector may find their face clearly at the entrance. Tracking then updates the region as they move. Behind the pillar, there may be no face pixels to observe. On the other side, the system must find the face again and decide how it relates to earlier appearances.

That last step is a common source of confusion. The person has not changed, but the available visual evidence has.

Why masks drift

A tracker estimates movement from the images and its previous state. When the picture changes quickly, it can follow the wrong pattern or lag behind.

The visible symptom is a mask that moves away from the face or stays where the face used to be.

More opaque coverage in that wrong place does not solve the problem. The location over time needs correction.

Why the face list can show duplicates

A person's appearance can change with angle, scale, lighting, and camera cuts. Grouping may separate those appearances into multiple entries.

Conversely, similar-looking or overlapping appearances can be associated incorrectly. The face list is a practical selection interface, not a verified identity system.

When hiding one person, inspect every relevant thumbnail and then check the complete output.

What Unseen's local presets do

Relaxed, Balanced, Thorough, and Maximum allocate different detection and tracking effort. Maximum performs frequent detection, while other modes use different intervals and tracking between observations.

Thorough is the recommended starting point in the interface. Maximum is useful to test when a face appears briefly or motion is difficult.

These are local settings. They do not tune cloud detection, and they do not guarantee correct tracking through a fully hidden interval.

Diagnose the visible failure

If no mask appears when the face first enters, investigate detection. If it starts correctly then wanders, investigate tracking. If it follows another person after a crossing, investigate association.

If the region stays in the right place but the face remains recognizable, evaluate the effect and surrounding identifiers.

Use these distinctions in a timestamp note before retrying. The missed-frame guide turns the diagnosis into practical next steps.

Try it on a short clip

Start with a difficult moment from your video, then inspect the exported result.

Open Unseen