Post-training Kimi K3 against our tikanga eval
You cannot post-train a model you do not have the weights to. That single fact decides most of this: if the goal is a model whose behaviour around Maori data is something we can change and be accountable for, it has to be a model we can actually train and host, not one we rent behind an API.
Why an open-weights model
It also decides where inference can run. A model we hold is a model that can sit on infrastructure we control, which is the difference between saying data stays in the country and being able to show it.
Why K3 specifically
K3 is open weights, strong enough to be worth the effort, and small enough that a full post-training run is affordable more than once. That last part matters more than it sounds: an eval is only useful if you can afford to fail it, change something, and run again.
The setup
The setup, the training data, the hyperparameters and the number of runs are recorded so the work can be repeated rather than taken on trust.
What we measured
We measured against the six axes described in the eval post, scored by hand, before and after. The comparison that matters is not against other models, it is against the same model before the run.
Results
RESULTS PLACEHOLDER: the per-axis before and after scores go here, with the cases that moved most and the cases that did not move at all. Nothing is written up until the numbers are in.
What this does and does not prove
What a result here would prove is narrow and worth stating plainly: that a specific model's behaviour on a specific set of cases changed in a specific direction. It would not prove the model is grounded in tikanga, and it would not make it safe to point at iwi data unsupervised. Those are claims no eval score earns.
Ready to put AI to work in your business?
We find the one workflow costing you the most time or the most leads, ship it into production, and prove what it saved. Businesses across New Zealand.