Post-training Kimi K3 against our tikanga eval

You cannot post-train a model you do not have the weights to. That single fact decides most of this: if the goal is a model whose behaviour around Maori data is something we can change and be accountable for, it has to be a model we can actually train and host, not one we rent behind an API.

6 September 2026 · 1 min read

Why an open-weights model

It also decides where inference can run. A model we hold is a model that can sit on infrastructure we control, which is the difference between saying data stays in the country and being able to show it.

Why K3 specifically

K3 is open weights, strong enough to be worth the effort, and small enough that a full post-training run is affordable more than once. That last part matters more than it sounds: an eval is only useful if you can afford to fail it, change something, and run again.

The setup

The setup, the training data, the hyperparameters and the number of runs are recorded so the work can be repeated rather than taken on trust.

What we measured

We measured against the six axes described in the eval post, scored by hand, before and after. The comparison that matters is not against other models, it is against the same model before the run.

Results

RESULTS PLACEHOLDER: the per-axis before and after scores go here, with the cases that moved most and the cases that did not move at all. Nothing is written up until the numbers are in.

What this does and does not prove

What a result here would prove is narrow and worth stating plainly: that a specific model's behaviour on a specific set of cases changed in a specific direction. It would not prove the model is grounded in tikanga, and it would not make it safe to point at iwi data unsupervised. Those are claims no eval score earns.

Ready to put AI to work in your business?

We find the one workflow costing you the most time or the most leads, ship it into production, and prove what it saved. Businesses across New Zealand.

Talk to an AI engineer