A public, reproducible safety benchmark for frontier AI models - focused on how they behave with children. Independent. Open source. Re-scored every quarter.
1,840
Scenario prompts
12
Languages tested
3
Independent graders
Quarterly
Re-evaluation
We evaluate every publicly available frontier chat model across 1,840 scenario prompts spanning four age-aware categories:
Every prompt is written from scratch by child-safety researchers and stress-tested across three age brackets - 5-7, 8-11, and 12-15. Each one has a documented expected behaviour and a machine-readable grading rubric.
We do not adapt adult red-team sets. The way a 9-year-old asks about self-harm, sex, drugs or violence looks nothing like a jailbreak.
Scenarios run in 12 languages - English, Spanish, Portuguese, French, German, Italian, Dutch, Polish, Arabic, Hindi, Mandarin, Japanese. Most kids using AI today don't type in English. The biggest safety regressions show up in non-English completions, so we score them explicitly.
Responses are graded by two independent LLM judges and one human reviewer. A score is only locked when at least 2 of 3 graders agree. Disagreements escalate to a senior child-safety reviewer.
Judges run on isolated infrastructure with no access to the model's identity. Every grader decision is logged so anyone can audit why a response passed or failed.
Models drift. A model that was safe on release can regress two weeks later after a silent provider update. So we re-score every model on the leaderboard every quarter using the latest public API endpoint. New models are added within two weeks of public release. Any score change greater than ±3 points triggers a public note in the changelog.
AstroSafe has no commercial relationship with the model makers we score. No sponsorship, no paid placement, no pre-publication review. The benchmark is funded entirely by the AstroSafe AI Safe Certificate programme - priced and disclosed separately.
When a model maker disagrees with a score, we publish their response alongside ours. We do not change scores under pressure.
All prompts, rubrics, grader code and raw model outputs are published on GitHub under an MIT licence. Anyone can clone the repo, point it at their own API keys and re-run the entire benchmark in under four hours on a single workstation.
If your numbers differ from ours by more than ±2 points on any category, open an issue - we'll dig in and publish a public post-mortem.
This benchmark is deliberately narrow. We do not score reasoning, code quality, multimodal performance, latency or cost. Plenty of great benchmarks already do. We measure one thing - how the model behaves when a child is on the other side of the conversation - and we measure it well.
No commercial relationship with the model makers we score. Funded by the certification programme.
Every scenario, prompt and grading rubric is public. Re-run the benchmark and verify the numbers yourself.
New models added on release. Existing models re-scored every quarter as providers patch behaviour.
Every prompt is written by child-safety researchers - not adapted from adult red-team sets.
Scenarios run in 12 languages so we catch safety gaps that only surface outside English.
The full benchmark harness lives on GitHub. Fork it, audit it, contribute new scenarios.
v1.0
Jan 28, 2026
Initial public release. 1,840 prompts, 12 languages, 9 model makers.
v0.9-beta
Nov 14, 2025
Closed beta with 4 model makers. Prompt set finalised, grader rubrics frozen.
v0.5-research
Aug 02, 2025
Internal research preview. Grader-agreement methodology validated.
Submit your model - or apply for the AstroSafe AI Safe Certificate and tell parents your product is built for their kids. We'll run the full benchmark, share a private read-out, and publish your score once you're ready.