Real-time face detection using YuNet and CLI automation
YuNet is a compact face detector available through OpenCV. A small CLI can use it to annotate local images and videos, with predictable batch outputs and checked failures. Detection finds face-like regions; it does not identify people or provide an authentication decision.
Understand YuNet’s strengths
YuNet offers a lightweight CPU-oriented option for detecting faces and facial landmarks. Accuracy and speed depend on image size, face size, pose, lighting, threshold, and hardware. See the OpenCV Zoo model documentation for the model and its published evaluation; measure your own workload before making real-time or accuracy guarantees.
Set up YuNet quickly
Create a Python environment and install one OpenCV wheel. This CLI does not open GUI windows,
so it uses the headless package; do not also install opencv-python into the same environment.
python3 -m venv venv
. venv/bin/activate
pip install opencv-python-headless==5.0.0.93
Download the tested model bytes, not a Git LFS pointer. This command pins the OpenCV Zoo revision:
curl --fail --location --output face_detection_yunet.onnx \
https://media.githubusercontent.com/media/opencv/opencv_zoo/47534e27c9851bb1128ccc0102f1145e27f23f98/models/face_detection_yunet/face_detection_yunet_2023mar.onnx
The model is 232,589 bytes with SHA-256
8f2383e4dd3cfbb4553ea8718107fc0423210dc964f9f4280604804ed2552fa4. Verify it with sha256sum on
Linux or shasum -a 256 on macOS. Keep the model beside the scripts. Use a compatible supported
Python version for the pinned wheel and keep the resulting environment reproducible.
Automate single-image detection
Save this shared implementation as face_tools.py. It uses a real resized frame for inference and
maps bounding boxes back to the original image, rather than changing the detector's size setting
without resizing its input.
from contextlib import contextmanager
import math
import os
from pathlib import Path
from tempfile import TemporaryDirectory
# Set before importing OpenCV so image decoding has an explicit pixel ceiling.
os.environ.setdefault('OPENCV_IO_MAX_IMAGE_PIXELS', '20000000')
import cv2
MODEL_PATH = Path(__file__).with_name('face_detection_yunet.onnx')
def create_detector():
return cv2.FaceDetectorYN.create(str(MODEL_PATH), '', (320, 320), 0.9, 0.3, 5000)
def annotate(detector, image, max_side=960):
height, width = image.shape[:2]
if width * height > 20_000_000:
raise ValueError('Frame exceeds the pixel limit')
scale = min(1.0, max_side / max(width, height))
small = cv2.resize(image, (max(1, round(width * scale)), max(1, round(height * scale))))
detector.setInputSize((small.shape[1], small.shape[0]))
_, faces = detector.detect(small)
count = 0 if faces is None else len(faces)
if faces is not None:
scale_x = width / small.shape[1]
scale_y = height / small.shape[0]
for face in faces:
x, y, w, h = face[:4]
left, top = max(0, round(x * scale_x)), max(0, round(y * scale_y))
right = min(width - 1, round((x + w) * scale_x))
bottom = min(height - 1, round((y + h) * scale_y))
cv2.rectangle(image, (left, top), (right, bottom), (0, 255, 0), 2)
return image, count
@contextmanager
def new_output(path):
output = Path(path).absolute()
if output.exists() or output.is_symlink():
raise ValueError('Output must not exist')
with TemporaryDirectory(prefix='.face-output-', dir=output.parent) as temporary:
candidate = Path(temporary) / ('result' + output.suffix)
yield candidate
if not candidate.is_file() or candidate.stat().st_size == 0:
raise ValueError('No nonempty output produced')
os.link(candidate, output)
def detect_image(detector, source, output):
source = Path(source)
if not source.is_file() or source.stat().st_size > 50 * 1024 * 1024:
raise ValueError('Use a regular image file no larger than 50 MiB')
image = cv2.imread(str(source))
if image is None:
raise ValueError('Cannot decode image')
annotated, count = annotate(detector, image)
with new_output(output) as candidate:
if not cv2.imwrite(str(candidate), annotated):
raise ValueError('Cannot encode output image')
return count
def detect_video(detector, source, output):
if not Path(source).is_file():
raise ValueError('Use a local video file')
capture = cv2.VideoCapture(str(Path(source).absolute()))
writer = None
frames = 0
detections = 0
try:
if not capture.isOpened():
raise ValueError('Cannot open video')
fps = capture.get(cv2.CAP_PROP_FPS)
width = int(capture.get(cv2.CAP_PROP_FRAME_WIDTH))
height = int(capture.get(cv2.CAP_PROP_FRAME_HEIGHT))
if not math.isfinite(fps) or fps <= 0 or width <= 0 or height <= 0:
raise ValueError('Invalid video timing or dimensions')
if width * height > 20_000_000 or width % 2 or height % 2:
raise ValueError('Use bounded, even video dimensions')
with new_output(output) as candidate:
writer = cv2.VideoWriter(str(candidate), cv2.VideoWriter_fourcc(*'mp4v'), fps, (width, height))
if not writer.isOpened():
raise ValueError('Cannot open video encoder')
while True:
ok, frame = capture.read()
if not ok:
break
if frame.shape[:2] != (height, width):
raise ValueError('Frame dimensions changed')
annotated, count = annotate(detector, frame)
writer.write(annotated)
frames += 1
detections += count
writer.release()
writer = None
if frames == 0:
raise ValueError('No frames decoded')
return frames, detections
finally:
capture.release()
if writer is not None:
writer.release()
Save this caller as detect_faces.py in the same directory:
import argparse
import sys
from face_tools import create_detector, detect_image, detect_video
import cv2
parser = argparse.ArgumentParser()
parser.add_argument('mode', choices=['image', 'video'])
parser.add_argument('input')
parser.add_argument('output')
args = parser.parse_args()
try:
detector = create_detector()
if args.mode == 'image':
print(f'Detected {detect_image(detector, args.input, args.output)} faces')
else:
frames, detections = detect_video(detector, args.input, args.output)
print(f'Processed {frames} frames with {detections} face detections')
except (OSError, ValueError, cv2.error):
sys.exit('Face detection failed; check input, model, codec, and destination')
Run python detect_faces.py image photo.jpg annotated.jpg. A valid image with no detected faces
still produces an output and reports zero detections. The output filesystem must support hard
links; existing files are never replaced by publication.
Batch-process folders
Reuse the same detector sequentially rather than sending an unpicklable lambda to a process pool or
letting every input overwrite output.jpg. Save this as batch_detect.py:
import argparse
from pathlib import Path
import sys
from face_tools import create_detector, detect_image
import cv2
parser = argparse.ArgumentParser()
parser.add_argument('input_directory')
parser.add_argument('output_directory')
args = parser.parse_args()
try:
source = Path(args.input_directory)
images = sorted(path for path in source.iterdir() if path.is_file() and not path.is_symlink()
and path.suffix.lower() in {'.jpg', '.jpeg', '.png'})
if not images:
sys.exit('No supported images found')
detector = create_detector()
output = Path(args.output_directory)
output.mkdir(mode=0o700)
failures = 0
for image in images:
try:
count = detect_image(detector, image, output / (image.name + '.jpg'))
print(f'{image.name}: {count} faces')
except (OSError, ValueError, cv2.error):
failures += 1
print(f'Failed: {image.name}', file=sys.stderr)
sys.exit(1 if failures else 0)
except (OSError, ValueError, cv2.error):
sys.exit('Cannot prepare batch; check model and directories')
Run python batch_detect.py images new-results. Full source filenames are retained before adding
.jpg, so photo.jpg and photo.png get separate outputs. Failed files count toward a nonzero
batch status; successful results remain available.
Handle video streams
Run python detect_faces.py video input.mp4 annotated.mp4 for a local constant-frame-rate video.
The output contains annotated video only: audio is not copied. The detection total counts
observations across frames, not distinct people. Variable-frame-rate timing and live camera
reconnection need a different workflow.
OpenCV's failed frame read can mean end-of-file or a decoder problem, and VideoWriter.write does
not report per-frame success. Check resulting frame counts, duration, and playback before
publishing. The example checks initialization and nonempty output, not complete media integrity.
Tune for real-time speed
Reduce max_side in annotate to actually resize inference frames, then measure the tradeoff on
small and distant faces. Keep one detector per worker process; do not share it across simultaneous
calls. A CPU wheel does not automatically provide CUDA support. GPU acceleration requires a
compatible OpenCV build, hardware, and a separately validated backend configuration.
For untrusted media, use isolated workers and whole-job time, memory, and disk limits. The image byte and pixel checks do not bound every video-decoder allocation or total processing time.
Compare alternative detectors briefly
| Detector | Useful starting point | Important limitation |
|---|---|---|
| YuNet | Compact face detection through OpenCV | Accuracy depends on scale, pose, and threshold |
| OpenCV DNN face models | A different pretrained model or backend | Verify the exact model and preprocessing |
| Haar cascades | Legacy lightweight detection workflows | Often less robust to pose and lighting variation |
Compare candidates on the same representative data instead of treating qualitative speed claims as a benchmark.
Wrap-up
A shared detector helper keeps image, folder, and video workflows consistent while preserving outputs and exposing failures. Verify the pinned model and actual annotations, then measure the performance and accuracy required by your application. For managed vision workflows, see Transloadit's Artificial Intelligence service.
