We built an eval for how well a model is grounded in tikanga
Ask a frontier model to define kaitiakitanga and it will answer well. Ask it to handle a request that quietly breaches kaitiakitanga, and it will usually comply, because nothing in its training made refusal the expected behaviour. The gap between those two responses is the thing we wanted to measure.
Knowing about tikanga is not the same as being grounded in it
Knowledge benchmarks reward recall. A model can score highly on questions about te ao Maori while still agreeing to repurpose data collected for one iwi programme into another, or summarising korero tuku iho for an audience it was never meant for. Recall and conduct are separate capabilities and the second one is the one that matters in production.
Why the existing benchmarks miss it
Every general-purpose benchmark we looked at tests knowledge in English, about Maori, from the outside. None of them test what a model does when it is put in a position where the tikanga-correct answer costs the user something. That is the entire problem, and it is not measurable by multiple choice.
The six axes we score against
The eval scores against the six principles of Maori data sovereignty published by Te Mana Raraunga: rangatiratanga, whakapapa, whanaungatanga, kotahitanga, manaakitanga and kaitiakitanga. We use those rather than inventing our own axes, because a framework that already carries authority is worth more than one that is convenient for us.
How a case is written
Each case is a realistic request with a tikanga-relevant tension inside it, written so that the unhelpful-looking answer is the correct one. A case is not a trick question. It is the kind of request an organisation actually sends: merge these two registers, summarise this hui, generate a mihi for this event, use last year's consultation data for this year's proposal.
Scoring, and who does it
Scoring is not automated. Each response is graded by a person against the principle it was written to test, with the reasoning recorded alongside the score. An LLM judge would be cheaper and would reproduce exactly the blind spot the eval exists to find.
What we are publishing, and what we are not
We are publishing the method, the axes and the case structure. We are not publishing the case set itself. An eval that is in the training data stops being an eval, and this one is small enough that leaking it would end it.
The results for the models we have run so far are being written up separately, alongside the post-training work. This post is the method.
Ready to put AI to work in your business?
We find the one workflow costing you the most time or the most leads, ship it into production, and prove what it saved. Businesses across New Zealand.