AI and Children's Data

AgentChamp - AI and Children's Data

Effective 25 August 2026


The short answer

No child's data is ever sent to an AI model. No child's data is ever used to train, fine-tune or evaluate one. Not ours, not a provider's, not in aggregate, not anonymised, not ever.

AI writes stories before children read them, from parameters an adult chose. It is a printing press, not an observer.

This document exists because "we use AI responsibly" is a claim anyone can make. The ICO's June 2026 *Edtech examined* audit found that most providers it examined were reusing children's data for internal purposes - product development, analytics, anonymisation, AI - without a lawful basis and without saying so, and that nearly 70% had become controllers of children's data without realising it. The detail below is so you can check we are not among them.


Where AI is used
1. Writing stories

An adult - a teacher, a parent, or our own content team - chooses a theme, a reading level and a style. A pipeline of AI models then researches, drafts, critiques and revises a story, and generates its cover artwork.

Models used: Google Gemini and Imagen, and Anthropic Claude, all accessed through Google Cloud Vertex AI in a Google Cloud project we control.

What the request contains: the content parameters. A theme like "a lighthouse in a storm". A year group. A style note. That is all - the job payload has no field for a user identifier, and no code path puts one there.

What happens before a child sees it: every generated story passes automated validation - reading level, word counts, page counts, vocabulary coverage, content safety - and a person can withdraw any story at any time.

2. Reading stories aloud

Story narration, quiz questions and dictionary words are turned into speech by ElevenLabs, Microsoft Azure Speech and Google Cloud Text-to-Speech.

These receive the published text and nothing else. No name, no account reference, no session identifier.

And the audio is synthesised once, when the story is published, and the same file is served to every reader. This matters more than it sounds. Synthesising per child would have been simpler to build and cheaper to run - and it would have meant sending a third party a request tied to a specific reader, every time a child pressed play. Generating once means a speech provider can never learn that a particular child read a particular book, because the request happens before any child has opened it.

3. Where AI is *not* used
  • Quizzes are not marked by AI. Questions and their answer keys are authored by people. Marking is a server-side comparison against the key. An earlier design would have generated questions with a language model; it was never enabled and the code path is dead.
  • Dictionary definitions are not generated at read time. They are written in-house, stored in our own database, and served from it. No network call leaves our infrastructure when a child taps a word.
  • Reading-level placement is not machine learning. It is a rules-based calculation over a child's reading history - explainable, inspectable, and overridable by any teacher.
  • There is no chatbot, no AI tutor, and no conversational surface of any kind in AgentChamp. A child cannot talk to a model, because there is nothing to talk to.

The four questions worth asking a supplier

Is children's data sent to an AI model? No. Story generation runs before any child is involved, from parameters an adult supplied. Speech synthesis receives published text. There is no runtime path from a child's activity to any model.

Is children's data used to train, fine-tune or evaluate models? No. Not by us, and not by our providers - the models are consumed through Google Cloud Vertex AI under enterprise terms that do not permit customer data to train foundation models, and in any case we send no customer data. For schools, this is a contractual obligation in the Data Processing Agreement, not a policy we could quietly change.

Is children's data anonymised and reused? No. Anonymising children's reading data for product development or resale is one of the specific practices the ICO flagged, and it is a practice that turns a processor into an undeclared controller. We do not do it. Our product analytics are counts of stories and quiz pass rates computed in our own database for teachers to read - not a dataset we mine or sell.

Who is the controller for the AI processing? Nobody's personal data is processed by it, so the question does not arise. We have assessed it anyway, because that is where providers most often slip: if we ever did send children's data to a model for our own purposes, we would become a controller for it and would have to say so and find a lawful basis. We have not, and we would tell you before we did.


Automated decisions about children

One thing in AgentChamp decides something about a child: the reading-level recommendation, which suggests books at a level of challenge suited to that reader.

Under Articles 22A–22D of the UK GDPR, as amended by the Data (Use and Access) Act 2025, significant automated decisions require safeguards. This one is not significant - but here are the safeguards anyway, because a school should not have to take that judgement on trust:

  • A person can always see it, and always override it. Teachers and parents can pin a child's level, or set a floor or a ceiling, directly.
  • It only changes suggestions. It does not restrict access to any book, feed into any assessment, or appear on any report card. A child can read anything in the library.
  • It is explainable. It is a rules-based calculation over reading history, not an opaque model, and its inputs are visible in the teacher dashboard.
  • It has no legal or similarly significant effect. Nothing follows from it except which books appear first.
  • It is the only one. We do not profile children for any other purpose.

Where a child has the optional additional-reading-support marker, the recommendation may offer books below the usual floor for their year group - which is the whole point of recording it.


Content safety

AI-generated stories reach children, so the safety question is real:

  • Every story is validated before publication - reading level, structure, and content appropriateness for the age band
  • Story generation prompts are constrained to age-appropriate themes, and a teacher's or parent's free-text style notes pass through a scanner before use
  • Adults, not children, initiate story creation. Children cannot generate anything
  • Any story can be withdrawn immediately, and a teacher who finds a problem can report it and have it pulled

We do not claim an automated pipeline is infallible. We claim it is bounded, reviewed, reversible, and that a person is accountable for what it produces.


Changes

If we ever wanted to use children's data with an AI model, that would be a material change: it would need a lawful basis, a new DPIA, a change to the DPA, notice to every school, and it would be published here first. We have no plans to do so.

Questions: privacy@agentchamp.co.uk.

AgentChamp keeps you signed in and remembers your reading settings using cookies and browser storage. What we store, in full.