ChatGPT hallucinates about subject-specific pedagogy.
A generic language model sounds convincing, even when its advice on testing or assessment isn't based on anything.
Case · Our own product · Education
An AI assistant (the engine behind TeachFoundry) that grounds teachers in formative assessment practice, with a source citation on every answer.

The challenge
ChatGPT and similar chatbots give convincing-sounding advice on testing and assessment. Even when that advice isn't based on anything. For subject-specific pedagogy that's a risk: wrong advice hits the student directly.
Teachers who take formative assessment practice seriously don't want loose tips. They want to know what an approach is based on. A recognisable framework, not guesswork from a language model.
There was no knowledge assistant curated specifically for formative assessment practice, with source citation on every answer. That's what DocentRAG adds.
What we built
This is our own product, known publicly as TeachFoundry. DocentRAG is the internal working title for the engine underneath. Together with a subject expert in formative assessment practice, we built a RAG knowledge assistant: teachers ask a question in plain language and get an answer from a curated knowledge base, with its source attached.
The system also generates ready-to-use lesson materials (lesson plans, exit tickets, rubrics, implementation plans) that export straight to PDF or Google Docs.
Technically this is called retrieval-augmented generation: the system first searches its own knowledge base for relevant fragments and only then lets the language model formulate an answer, based on what it found. Not on whatever it happens to “know”. That way, every answer stays traceable to a concrete document from the formative assessment practice framework.
The sheet the answer rests on, in the knowledge base.Source · Core standards 2024, §4.2
No sourceLow confidence · no source in the knowledge base
How it works
In plain language, the way you'd ask a colleague. About formative assessment practice, testing or grading.
Not the open internet and not the language model's memory: only the walled-off knowledge base, based on recognised literature.
Summary, elaboration, sources and follow-up questions. Plus an indicator: high, medium or low confidence.
From lesson plan to exit ticket, rubric or implementation plan, ready to export straight to PDF or Google Docs.
The system supplies the grounding and the material, the teacher knows the class and decides what goes into the lesson.
Who it's for
Teachers and schools use DocentRAG (publicly: TeachFoundry) to ask questions about formative assessment practice and quickly get ready-to-use lesson material, grounded in recognised literature.
Why
A generic language model sounds convincing, even when its advice on testing or assessment isn't based on anything.
Advice without reference to a pedagogical framework. You can't check where it comes from, or whether it matches what's already known about formative assessment practice.
A chatbot hands back a block of text, not a lesson plan, rubric or exit ticket you can actually use in class tomorrow.
The same question about formative assessment practice gets put to a chatbot separately in a hundred places, with no shared, curated knowledge base behind it.
The difference
A general-purpose chatbot often sounds convincing on educational questions, even when the answer is a hallucination. The model fills a gap in its knowledge with something that sounds plausible about subject-specific pedagogy, without showing it.
DocentRAG works the other way round: the answer only comes after the system has found something in the curated knowledge base. No match (or a low confidence score) and you see that reflected in the answer straight away.
An ordinary chatbot sounds convincing. DocentRAG shows where the answer comes from, and how sure it is.
How we approached it
Around 196 documents per language instance, put together with a subject expert in formative assessment practice. Based on recognised literature, including the Toetsrevolutie framework (René Kneyber) and the HOP model.
No answer without a locatable source. Structure: summary → elaboration → sources → follow-up questions.
Every answer shows high, medium or low confidence. Questions outside the knowledge domain are refused rather than have the system make something up.
Separate, live instances in Dutch, English and for Hong Kong. Each with its own language and its own knowledge base. The English one goes to market under its own brand, on the same engine.
The system supplies the grounding and the material. The teacher judges and decides what goes into the classroom.
What it means for teachers
No hunting through books or loose articles: the question in plain language, the answer with its source attached.
Lesson plans, exit tickets, rubrics and implementation plans. Ready to export, instead of drafting them yourself.
Every answer is rooted in a coherent framework for formative assessment practice, not in loose tips from a generic chatbot.
The system supplies, the teacher weighs up what fits the class. That stays human work.
In numbers
per language instance
Documents in the curated knowledge base
live
Language instances: Dutch, English, Hong Kong
types
Material types the system generates ready-to-use
Want to look for yourself first? The supplier scan and the Chain check are free and need no conversation.
The supplier scanIn a half-hour conversation we'll see whether your knowledge domain lends itself to the same thing.
Book an intro callThree different worlds, one principle: the AI does the work, the human checks and approves.
Read the case