NEWVectors or files. Pick a path.Start →
    Search & Discovery
    7 min read
    Updated 2026-10-09

    How Do I Keep Footage of Minors Out of a Video Training Dataset?

    To keep footage of minors out of a video training dataset, screen at ingest: detect and group faces across cuts, estimate each person's age, and send anyone near the cutoff or too unclear to age to human review. On our 650-face evaluation, a cutoff of 18 caught 82% of minors.

    Training Data
    Content Moderation
    Age Estimation
    Face Detection
    Video

    How do I keep footage of minors out of a video training dataset?



    Screen the corpus at ingest, before it is copied into training. Cut each video into segments, sample keyframes, detect every face, group the looks at the same person across cuts, and estimate each person's age. Send anyone estimated near or under your cutoff to human review, along with every face too small or unclear to age, and keep only segments with no such person in the cleared set. Record the decision and its reason for every segment, so you can show a regulator, a customer or a court what was excluded and why.

    Why can't I just set the age cutoff at 18?



    Age estimation from a face has an error of several years, so a cutoff at the legal age lets through a share of minors who look older. Raising the cutoff catches more of them and sends more adults to review. Measured on our own evaluation set of 650 labelled faces (320 minors, 330 adults), one frame each:

    CutoffMinors caughtAdults sent to review
    1882%4.8%
    30100%94.5%
    Neither end is free. At 18 one minor in five is missed; at 30 nearly all adult footage needs a person to look at it. Pick the point on that curve your review team can staff, and decide it in writing before the first run.

    What about faces the model cannot age?



    Treat them as unknown and route them to review. In the same evaluation, 39% of detections went to review as unaged: no usable face, a face under 24 pixels, or a detector confidence under 0.70. None of them should be passed as adult.

    Gate the face detector before you age anything. A confidence threshold of 0.70 cleared six phantom "minors" the detector had found in things that were not faces, while keeping 85% of genuine faces. Without that gate, false faces inflate the review queue and train reviewers to wave items through.

    Why group faces across cuts?



    The same person appears many times in a video, from different angles and distances. Estimate age per person over several looks, and decide per person, so one bad frame does not clear someone and one good frame does not condemn a whole segment. A person whose estimates fall within a few years of the cutoff goes to review with all their looks attached.

    What should the screening record?



  1. Per segment: the decision, the reason, and counts of people, minors, unaged faces and borderline faces.
  2. A review queue ranked by risk, with the evidence for each item attached.
  3. The cleared set as a filterable export for the training job.


  4. Mixpeek's design notes list "no trace of why a result came back" among the failures that sink systems like this. A screen whose exclusions carry no reasons cannot be audited, and an exclusion log is often the thing you are asked for. A cleared segment means eligible under your screening policy. Rights and legal clearance are separate checks.

    How do I do this with Mixpeek?



    Mixpeek reads your files where they already live, in S3, GCS or any bucket, and turns what's in them into data your software can search, classify and moderate. The video moderation template deploys the extraction layer in one click: segments with lineage to their source video, keyframes per segment, plain-text segment search, and a brand-mark reference index with its matching retriever. The measurements above come from the evaluation behind that template.

    What it does not deploy on Mixpeek Cloud is stated on the template page. Face detection and public-figure matching run on a dedicated or customer-hosted deployment. The age estimation extractor is not published yet, and the verdict roll-ups, review queue and cleared-set export are not available yet. Teams preparing training data with face screening today run it on a dedicated deployment. Related: best AI content moderation tools and why video search misses on-screen text.

    Frequently Asked Questions



    How do I keep footage of minors out of a video training dataset?



    Screen at ingest: segment the video, detect and group faces across cuts, estimate each person's age, and send anyone near or under your cutoff, plus every face too unclear to age, to human review. Keep only segments with no such person, and record the reason for every exclusion.

    How accurate is AI age estimation for screening minors?



    Use it to rank faces for review and leave borderline decisions to a person. On our evaluation of 650 labelled faces, a cutoff of 18 caught 82% of minors; catching all of them took a cutoff of 30, which sent 94.5% of adult footage to review.

    What should happen to faces that are too small or blurry to age?



    Send them to review as unaged. In our evaluation that was 39% of detections, mostly faces under 24 pixels or below a 0.70 detector confidence.

    Does a cleared dataset mean the footage is legally safe to train on?



    No. Cleared means the segment passed your screening policy. Rights, consent and licensing are separate checks.
    Managed Mixpeek

    Put multimodal search to work

    Connect a bucket and Mixpeek runs the whole multimodal search pipeline for you: extraction, indexing, and search over your own objects. No models to wire up, nothing to host.

    Start with Managed
    MVS · bring your own

    Already have vectors?

    Keep your embeddings on your own cloud and run dense, sparse, and BM25 search directly on object storage. From $25/mo.

    Start with MVS

    Run this on your own data

    Point Mixpeek at the storage you already have and search your video, images, audio, and documents the way this guide describes. Build starts at $25/mo for up to 1M vectors.

    Search your own archiveRead Docs