Up-to-date results for every frontier model, with open methodology. We benchmark every model Aegis routes to - so the models powering your product are vetted before they reach a child.
34
Models scored
9
Model makers
1,840
Scenario prompts
Quarterly
Re-evaluation
Independent, reproducible, kid-focused - and re-scored every quarter so the numbers stay honest.
We have no commercial relationship with the model makers we score. Funded by the certification programme.
Every scenario, prompt and grading rubric is public. Re-run the benchmark and verify the numbers yourself.
New models added on release. Existing models re-scored every quarter as providers patch behaviour.
Every prompt is written by child-safety researchers - not adapted from adult red-team sets - so scores reflect what kids actually encounter.
Scenarios run in 12 languages so we catch safety gaps that only surface outside English - where most kids actually talk to AI.
The full benchmark harness lives on GitHub. Fork it, audit it, contribute new scenarios - the methodology evolves in the open.
Submit your model - or apply for the AstroSafe AI Safe Certificate and tell parents your product is built for their kids. We'll run the full benchmark, share a private read-out, and publish your score once you're ready.