How we test for minors and explicit content
We ran nine systems for detecting minors and seven systems for detecting explicit content against the same 394 images from our own platform. No single system won both jobs. The set has since grown to 1,020 images for the prompt work further down. This page names every system we tested, shows the numbers, and shows why layering them beats any one of them.
Benchmark run: 2026-09-02
Why we measured
A content filter has one job. Catch what should never appear, and let everything else through. Most classifiers are good at one side of that job and weak at the other. A precise filter misses real cases. A complete filter blocks real users.
We wanted an actual number for that trade-off instead of a guess. So we built one test set and ran every system we could get access to against it, side by side, on our own images.
What we measured
We tested 394 images. Some are real photographs from a public face dataset with published age labels. Some are images generated on our own platform. Every image was checked by an independent judge model before it went into the set, and no image in the set is a sexual depiction of a minor.
| Category | Images | Note |
|---|---|---|
| Generated images of minors | 69 | Labeled by an independent judge model |
| Real photos of minors | 60 | FairFace validation set, dataset age labels |
| Adults | 233 | |
| No people | 32 | |
| Total | 394 |
Two smaller groups sit inside the adult count above, not on top of it: 40 hard-negative adults (young-looking adults, the hardest false-positive test) and 63 adult explicit-content positives (flagged by a classifier sweep, confirmed by the judge).
How each system performed: minor detection
Nine systems, one shared test set. “Recall” is the share of real minors a system caught. “Precision” is the share of its blocks that were actually correct. A system can look strong on one and weak on the other. Each row uses that system’s own default cutoff, which is why the open-weight model reads conservative here and stronger in the layered results further down, where it runs at a tuned operating point.
| System | Precision | Recall | F1 | ROC-AUC | False positives (adults) | Speed |
|---|---|---|---|---|---|---|
| Grok 4.6 vision judge | 93.4% | 88.4% | 0.908 | 0.925 | 3.4% | 5.3 s |
| GPT-5.5 vision judge | 89.9% | 83.0% | 0.863 | 0.879 | 7.1% | 2.6 s |
| MiVOLO v2 age model | 98.8% | 64.3% | 0.779 | 0.820 | 0.4% | 67 ms |
| Shieldstral 3B, two queries | 100.0% | 58.1% | 0.735 | 0.958 | 0.0% | 1.4 s |
| Face detector + age classifier, tuned | 98.5% | 51.2% | 0.673 | 0.903 | 0.4% | 162 ms |
| Face detector + age classifier, raw | 97.1% | 51.9% | 0.677 | 0.903 | 0.9% | 180 ms |
| Age classifier + Grok clearance | 98.5% | 50.4% | 0.667 | 0.903 | 0.4% | 175 ms |
| Hosted explicit-content filter model | 78.1% | 38.8% | 0.518 | 0.664 | 6.0% | 2.3 s |
| Shieldstral 3B, one query | 100.0% | 38.0% | 0.551 | 0.956 | 0.0% | 1.4 s |
| OpenAI moderation endpoint | N/A | 0.0% | N/A | 0.500 | 0.0% | 364 ms |
OpenAI’s moderation endpoint isn’t plotted here: it flagged zero minors on this set, so its precision is undefined. It is still in the table below.
How each system performed: explicit content
Seven systems, scored on the same images for adult explicit content. Most of the positives in our set are covered or suggestive images rather than exposed nudity, which is a harder test than it sounds.
| System | Precision | Recall | F1 | ROC-AUC | Speed |
|---|---|---|---|---|---|
| Falcons AI image classifier (self-hosted) | 98.3% | 92.1% | 0.951 | 0.997 | 47 ms |
| Grok 4.6 vision judge | 89.7% | 96.8% | 0.931 | 0.974 | 5.3 s |
| Hosted explicit-content filter model | 56.1% | 73.0% | 0.634 | 0.811 | 2.3 s |
| Shieldstral 3B, two queries | 100.0% | 28.6% | 0.444 | 0.970 | 1.4 s |
| Shieldstral 3B, one query | 100.0% | 28.6% | 0.444 | 0.970 | 1.4 s |
| OpenAI moderation endpoint | 100.0% | 20.6% | 0.342 | 0.994 | 364 ms |
| NudeNet v3 detector | 83.3% | 7.9% | 0.145 | 0.592 | 258 ms |
Why every single approach fails somewhere
- The most precise system we tested missed about half the minors in our set. Precision alone is not a safety story.
- The most complete systems are slow and cost money. Grok took about five seconds per image. GPT-5.5 took about two and a half seconds. Neither can sit in front of every upload.
- One hosted filter model was the weakest system we measured on both minor detection and explicit content. A model that is mediocre at two jobs loses to two models that each do one job well.
- We tried eight versions of the same open-weight model’s prompt. Small wording changes moved which images got flagged, but none of them made the model meaningfully better at telling minors from adults in the first place. A better prompt moves the operating point. It does not move the curve.
Prompt variants we tried
All eight variants are the same open-weight model with a different question or instruction. ROC-AUC measures how well a system can separate minors from adults across every possible cutoff, independent of where that cutoff is set. Recall and false-positive figures below are shown at two fixed reference points so the variants can be compared; the reference points are not operating settings.
| Prompt variant | ROC-AUC | Recall, reference point A | False positives, point A | Recall, reference point B | False positives, point B |
|---|---|---|---|---|---|
| One query (under 18) | 0.956 | 84.5% | 6.9% | 45.0% | 0.0% |
| Two queries (under 18, child) | 0.958 | 85.3% | 6.9% | 59.7% | 0.0% |
| Plus a teenager query | 0.946 | 100.0% | 81.5% | 99.2% | 52.8% |
| “Anyone in the image” wording | 0.954 | 69.0% | 2.6% | 58.1% | 0.0% |
| Plus an inverted adult query | 0.940 | 100.0% | 100.0% | 100.0% | 92.3% |
| Instruction: judge by proportions | 0.955 | 88.4% | 11.2% | 59.7% | 0.0% |
| Instruction: answer yes when unsure | 0.958 | 75.2% | 4.3% | 59.7% | 0.0% |
| Face crops plus whole image | 0.939 | 89.1% | 23.6% | 64.3% | 1.3% |
Why layers win
No single system solved this on its own. A cascade did: run the fast open-weight model on every image, then send only the cases it is unsure about to a slower AI judge for a second opinion. The result keeps false positives close to the most precise single system while catching far more real minors.
| Setup | Recall | False positives (adults) | Precision |
|---|---|---|---|
| Face detector + age classifier | 51.2% | 0.4% | 98.5% |
| Shieldstral alone | 85.3% | 6.9% | 87.3% |
| Grok alone | 88.4% | 3.4% | 93.4% |
| Shieldstral, ambiguous cases resolved by Grok | 80.6% | 1.3% | 97.2% |
The judge’s prompt is part of the layer
A second opinion that can only overturn a block can afford to lean toward “minor” when it is unsure; a wrong lean costs nothing, the block was already there. Put the same judge in charge of the ambiguous cases in a cascade and that lean starts creating blocks. We measured the same judge with both prompts on the same set. The deciding prompt gives up recall on faces every model places between 16 and 22, and gains a false-positive rate of zero. Neither prompt is right on its own; the job the layer has decides which one fits.
| Judge prompt | Recall, alone | False positives, alone | Recall, in cascade | False positives, in cascade | Precision, in cascade |
|---|---|---|---|---|---|
| Prompt written to clear blocks (lean minor when unsure) | 88.4% | 3.4% | 80.6% | 1.3% | 97.2% |
| Prompt written to decide (no lean, judge by proportions) | 66.7% | 0.0% | 69.0% | 0.0% | 100.0% |
A bigger boundary set
Prompt questions like that one live or die on faces near the age-18 line, and 394 images did not carry enough of them. So the set grew to 1020 items, all of them real photographs with an exact published age in the disputed band: 300 more faces from a public face dataset, and 323 whole-body person crops in real scenes with exact ages from 13 to 25, balanced on either side of 18.
| Split | Images | Note |
|---|---|---|
| FairFace validation, ages 10-19 | 182 | Real photographs, CC BY 4.0 |
| FairFace validation, ages 20-29 | 165 | Real photographs, CC BY 4.0 |
| LAGENDA person crops, ages 13-17 | 161 | Whole-body crops in real scenes, CC BY 2.0 |
| LAGENDA person crops, ages 18-25 | 162 | Whole-body crops in real scenes, CC BY 2.0 |
Two things to hold in mind reading the numbers below. The 10-19 bracket of the face dataset contains 18 and 19 year olds, so its recall column is a band rather than a ceiling. And the 18-to-25 split is the hardest adult set anyone here has measured on: every judge in this test, including the commercial ones, reads a large share of real 18 to 25 year olds as possibly under 18.
Asking for a number beats asking for a category
Six prompts for the same open-weight judge, Qwen3-VL-8B, on the same 1020 images. Five ask for a category, minor or adult, in different words. One asks for a single number, the apparent age in years, and treats under 18 as a minor.
| Prompt variant | Faces 10-19, recall | Faces 20-29, false positives | Crops 13-17, recall | Crops 18-25, false positives | Cascade recall | Cascade false positives | Cascade precision |
|---|---|---|---|---|---|---|---|
| Clearance-style prompt | 97.0% | 56.0% | 99.0% | 87.0% | 0.900 | 0.311 | 0.700 |
| Decide prompt | 87.0% | 25.0% | 96.0% | 65.0% | 0.870 | 0.245 | 0.740 |
| Decide, ties go to adult | 91.0% | 34.0% | 96.0% | 73.0% | 0.880 | 0.259 | 0.730 |
| Decide, proportions only | 88.0% | 27.0% | 96.0% | 64.0% | 0.880 | 0.237 | 0.750 |
| Two-step (describe, then decide) | 82.0% | 17.0% | 93.0% | 51.0% | 0.850 | 0.205 | 0.770 |
| Numeric apparent age, minor under 18 | 68.0% | 5.0% | 80.0% | 22.0% | 0.740 | 0.104 | 0.850 |
Rewording the categorical question did nothing. Telling the model to give ties to the adult raised false positives on every adult split instead of lowering them. Changing the answer schema to a number did the work: the numeric prompt cut false positives on real 18 to 25 year olds from 65% to 22% and on real 20 to 29 year olds from 25% to 5%. The bill is recall, and it is paid at the 16-to-18 edge rather than on children.
A number also turns the age line into a dial, which a yes or no answer cannot be. The same run scored at three cutoffs.
| Cutoff | Cascade recall | Cascade false positives | Cascade precision |
|---|---|---|---|
| Minor under 18 | 0.740 | 0.104 | 0.850 |
| Minor under 20 | 0.810 | 0.171 | 0.790 |
| Minor under 22 | 0.830 | 0.192 | 0.770 |
A larger version of the same open-weight judge, Qwen3-VL-32B, scored 90.0% recall, 3.0% false positives and 0.940 precision on the original 394 images. That is inside sample noise of the smaller model, at 3.4x the latency and a 48 GB card minimum.
The same prompt on Grok
A prompt result that only holds for one model is a quirk, not a finding. So the three prompts ran again on a 200-image sample of the same set, this time on a commercial judge as well: 60 lagenda, ages 13-17, 60 lagenda, ages 18-25, 40 fairface, ages 10-19 and 40 fairface, ages 20-29.
| Judge and prompt | Crops 13-17, recall | Crops 18-25, false positives | Faces 10-19, recall | Faces 20-29, false positives |
|---|---|---|---|---|
| Grok, numeric age, minor under 18 | 90.0% | 28.0% | 75.0% | 8.0% |
| Grok, decide prompt | 87.0% | 37.0% | 75.0% | 8.0% |
| Grok, clearance-style prompt | 97.0% | 82.0% | 88.0% | 35.0% |
| Qwen3-VL-8B, numeric age, minor under 18 | 82.0% | 22.0% | 60.0% | 10.0% |
| Qwen3-VL-8B, decide prompt | 95.0% | 73.0% | 88.0% | 30.0% |
For Grok the numeric prompt beat the categorical one on both axes at once, catching more real 13 to 17 year olds while flagging fewer real 18 to 25 year olds, and tied it everywhere else. The finding transfers across models. Asking a vision model for a number is a different question than asking it for a verdict, and the number is the better one.
Caveats
- Sample sizes are modest: 129 minors (69 generated, 60 real) and 233 adults for the minor task, 63 positives for explicit content. Recall figures carry roughly plus or minus 13 percentage points of uncertainty. Treat any gap under 10 points between two systems as noise, not a real difference.
- Real-photo labels come from a published face dataset with independent age labels. Generated-image labels come from an AI judge model, so that judge’s own score on the generated split is not a fully independent check. The real-photo results are the cleaner comparison.
- A generated image was never a real person at a real age, so age labels on that split are judgment calls rather than dataset ground truth. Several outright misses across systems are genuine disagreements between judges about images near the age-18 boundary, not clear errors.
- Two commercial vendors that offer a dedicated minor-detection class were not part of this benchmark. They were not available to us for this run.
- The two public sets in the boundary work are real photographs. Nothing public covers photoreal generated people with a known age, so the hardest case for a generation platform has no dataset behind it.