Methodology · v1

How the AstroSafe benchmark works.

A public, reproducible safety benchmark for frontier AI models - focused on how they behave with children. Independent. Open source. Re-scored every quarter.

1,840

Scenario prompts

12

Languages tested

3

Independent graders

Quarterly

Re-evaluation

The full methodology

Eight steps, in the open.

1

Scope

We evaluate every publicly available frontier chat model across 1,840 scenario prompts spanning four age-aware categories:

  • Refusal - does the model refuse harmful, age-inappropriate or grooming-style requests, however they're framed?
  • Tone - is the response age-appropriate for a 5-, 9- or 13-year-old?
  • Privacy - does it avoid asking for or retaining minors' personal information?
  • Honesty - no sycophancy, hallucination, or pretending to be a friend or clinician.
2

Prompts

Every prompt is written from scratch by child-safety researchers and stress-tested across three age brackets - 5-7, 8-11, and 12-15. Each one has a documented expected behaviour and a machine-readable grading rubric.

We do not adapt adult red-team sets. The way a 9-year-old asks about self-harm, sex, drugs or violence looks nothing like a jailbreak.

3

Languages

Scenarios run in 12 languages - English, Spanish, Portuguese, French, German, Italian, Dutch, Polish, Arabic, Hindi, Mandarin, Japanese. Most kids using AI today don't type in English. The biggest safety regressions show up in non-English completions, so we score them explicitly.

4

Grading

Responses are graded by two independent LLM judges and one human reviewer. A score is only locked when at least 2 of 3 graders agree. Disagreements escalate to a senior child-safety reviewer.

Judges run on isolated infrastructure with no access to the model's identity. Every grader decision is logged so anyone can audit why a response passed or failed.

5

Re-evaluation

Models drift. A model that was safe on release can regress two weeks later after a silent provider update. So we re-score every model on the leaderboard every quarter using the latest public API endpoint. New models are added within two weeks of public release. Any score change greater than ±3 points triggers a public note in the changelog.

6

Independence

AstroSafe has no commercial relationship with the model makers we score. No sponsorship, no paid placement, no pre-publication review. The benchmark is funded entirely by the AstroSafe AI Safe Certificate programme - priced and disclosed separately.

When a model maker disagrees with a score, we publish their response alongside ours. We do not change scores under pressure.

7

Reproducibility

All prompts, rubrics, grader code and raw model outputs are published on GitHub under an MIT licence. Anyone can clone the repo, point it at their own API keys and re-run the entire benchmark in under four hours on a single workstation.

If your numbers differ from ours by more than ±2 points on any category, open an issue - we'll dig in and publish a public post-mortem.

8

What we don't measure

This benchmark is deliberately narrow. We do not score reasoning, code quality, multimodal performance, latency or cost. Plenty of great benchmarks already do. We measure one thing - how the model behaves when a child is on the other side of the conversation - and we measure it well.

Principles

Six commitments behind every score.

Independent

No commercial relationship with the model makers we score. Funded by the certification programme.

Reproducible

Every scenario, prompt and grading rubric is public. Re-run the benchmark and verify the numbers yourself.

Up to date

New models added on release. Existing models re-scored every quarter as providers patch behaviour.

Kid-focused

Every prompt is written by child-safety researchers - not adapted from adult red-team sets.

Multilingual

Scenarios run in 12 languages so we catch safety gaps that only surface outside English.

Open source

The full benchmark harness lives on GitHub. Fork it, audit it, contribute new scenarios.

Versions

How the benchmark has evolved.

v1.0

Jan 28, 2026

Initial public release. 1,840 prompts, 12 languages, 9 model makers.

v0.9-beta

Nov 14, 2025

Closed beta with 4 model makers. Prompt set finalised, grader rubrics frozen.

v0.5-research

Aug 02, 2025

Internal research preview. Grader-agreement methodology validated.

Get listed

Want your model on the leaderboard?

Submit your model - or apply for the AstroSafe AI Safe Certificate and tell parents your product is built for their kids. We'll run the full benchmark, share a private read-out, and publish your score once you're ready.

Explore Aegis 2.0