Model Benchmark

The Model Benchmark for AI child safety

Up-to-date results for every frontier model, with open methodology. We benchmark every model Aegis routes to - so the models powering your product are vetted before they reach a child.

Coming soon

34

Models scored

9

Model makers

1,840

Scenario prompts

Quarterly

Re-evaluation

Methodology

Six principles behind every score.

Independent, reproducible, kid-focused - and re-scored every quarter so the numbers stay honest.

Independent

We have no commercial relationship with the model makers we score. Funded by the certification programme.

Reproducible

Every scenario, prompt and grading rubric is public. Re-run the benchmark and verify the numbers yourself.

Up to date

New models added on release. Existing models re-scored every quarter as providers patch behaviour.

Kid-focused

Every prompt is written by child-safety researchers - not adapted from adult red-team sets - so scores reflect what kids actually encounter.

Multilingual

Scenarios run in 12 languages so we catch safety gaps that only surface outside English - where most kids actually talk to AI.

Open source

The full benchmark harness lives on GitHub. Fork it, audit it, contribute new scenarios - the methodology evolves in the open.

Get listed

Want your model on the leaderboard?

Submit your model - or apply for the AstroSafe AI Safe Certificate and tell parents your product is built for their kids. We'll run the full benchmark, share a private read-out, and publish your score once you're ready.