Kaketa tried a vision model for checking handwritten kanji, then replaced it with a deterministic stroke matcher. How it works and how the two compare on my own attempts.
Overview
Kaketa is an app for practicing handwritten Japanese words with kanji. You see the reading, write the word in squares with Apple Pencil or a finger, show the answer with the stroke order of every character, and decide: got it, or not yet. The check is the part between showing the answer and deciding: something that looks at what was written and says whether it looks right, is close, or isn’t quite there, and ideally which stroke is off.
The first version asked a vision model. It sent a picture of the strokes next to the printed answer to OpenAI’s Decisions API and showed its answer. It had two faults: it said “correct” for almost everything, and when it didn’t, it couldn’t say which stroke it minded.
Kaketa has the strokes themselves, though, not just a picture of them: PencilKit gives each stroke as a path, and KanjiVG gives every character’s strokes in order. So the check can be geometry instead of a model. Kaketa now compares each stroke drawn with the standard one, on the device, and that comparison is the verdict. None of this is new: matching strokes against reference data is the standard way to grade handwriting practice, and this note is about why the model lost to it. OpenAI is gone from the app; a strictness setting took its place.
What the model did
To judge the check, Kaketa has an optional attempt log. With it on, every word written is kept with the picture, the strokes, the verdict shown, the decision (got it or not yet), and a thumbs up or down on the verdict shown. The decision is about the writing; the thumbs are about the check, a cleaner label. The numbers below are from my own testing, 124 attempts in a single day, and should be read as such: one hand (the developer’s, not a learner’s), one iPad, and almost all of the attempts were traced over the stroke guide, so nearly every one was decided “got it”.
OpenAI got a 320-point composite: the strokes on the left, the printed answer on the right, with the pen width normalized. It answered with probabilities for correct, close, and wrong, which Kaketa turned into a verdict with a display rule (not sure unless correct and close together reach 60 percent; close only when it leads). On 116 checked attempts it showed “correct” 107 times, “close” 6 times, and “wrong” 3 times. I endorsed the verdict I saw on 67 of 90 thumbs. The verdict agreed with my decision on 101 of 115.
The bigger problem wasn’t the agreement rate. The decision model can’t point at the wrong stroke, takes a few seconds, and needs the network.
Apple’s on-device Foundation Model was tried with the same picture for a while and did worse, agreeing with my thumbs on 11 of 35, so it isn’t covered here.
The stroke matcher
The matcher is about 300 lines of Swift with no dependencies. Its per-stroke checks are the ones Hanzi Writer uses in its quiz mode, which has graded strokes in the browser for years. The difference is that Hanzi Writer knows which stroke you are drawing, because it asks for them one at a time, while Kaketa grades a finished word, so it first has to work out which of your strokes is which.
- Each PencilKit stroke is sampled along its path, put in the square it starts in, and scaled into KanjiVG’s 109-point box. Every threshold below is in that box, so it means the same on an iPhone and an iPad.
- KanjiVG’s SVG paths for the character are flattened to polylines. Both sides are resampled to 32 points evenly spaced along their length.
- A cost matrix holds the average point-to-point distance between every drawn stroke and every reference stroke, taking the smaller of the stroke as drawn and reversed. The Hungarian method finds the assignment with the lowest total cost. A pair farther apart than 38 points isn’t a pair: the reference stroke counts as missing and the drawn one as extra.
- Each pair is then judged on Hanzi Writer’s five checks, and a pair that fails any of them is “off”, with the failed checks named for the developer build.
| Check | Passes when |
|---|---|
| Distance | The average distance between corresponding points is at most 19 points, in KanjiVG’s 109-point box: about a sixth of its width |
| Ends | The start is within 27 points of the reference start, and the end within 27 points of its end |
| Direction | Each drawn segment’s best cosine against the reference’s segments averages above zero |
| Shape | After centring both curves and scaling each to unit size, the discrete Fréchet distance between them is at most 0.4 (no unit, since both are scaled), with the reference also tried rotated by ±π/32 and ±π/16 |
| Length | The drawn length is at least 35 percent of the reference length, each length padded by 2.7 points so dots don’t divide by nothing |
Hanzi Writer’s defaults are in a 1024-point box: average distance 350 scaled by 0.5, start and end 250, Fréchet 0.4, length ratio 0.35 with a 25-point cushion. Kaketa scales them to 109.
The verdict for a character comes from counting problems, that is, strokes missing, extra, or off:
let problems = missing.count + failed + extra.count
if problems == 0 { return backwards > 0 ? .close : .correct }
if problems == 1 { return .close }
// A long character may have a few slips and still be close.
if passed >= Int((Double(total) * 0.6).rounded(.up)) && problems <= max(1, total / 3) { return .close }
return .wrong
A word’s verdict is its worst character’s. Under the verdict, the app prints a line per character, such as “日 4/4 · 本 4/5, 1 missing”, and the stroke numbers on the grid show which strokes it means. The whole thing runs in a few milliseconds when the answer is shown.
Backwards strokes
The first benchmark flagged the first stroke of 手 as off every time. I had been drawing it left to right; KanjiVG has it right to left. The fix: a stroke that fails as drawn but passes reversed counts as drawn backwards. It passes, the character is at most “close”, and the stroke’s number on the grid turns amber with a small arrow.
How they compare
Stroke data was added to the log partway through the day, so the matcher can be run only on the 57 attempts that have it. On those, every attempt was decided “got it” and OpenAI showed “correct” on all but one, so the decision column says little. The thumbs column is the one that matters: when I disagreed with “correct”, did the matcher also see a problem?
| Grader | Agrees with my thumbs | Matches my decision |
|---|---|---|
| Stroke matcher | 51 / 55 | 56 / 57* |
| OpenAI, as shown | 43 / 55 | 57 / 57 |
The 57 attempts with stroke data. A grader agrees with a thumbs up when it repeats the verdict I saw, and with a thumbs down when it departs from it. The matcher said correct 42 times, close 14, wrong once. *Every attempt was decided “got it”, including the 手 ones where I didn’t know I had a stroke backwards.
The matcher’s four disagreements were three 手 attempts where I accepted “correct” despite the backwards stroke, and one 学 where it found one of eight strokes off. The twelve times I disagreed with OpenAI’s “correct”, the matcher saw a problem on eleven.
Since the matcher became the verdict, the first 13 live attempts have had 12 thumbs, 11 of them up. The one down was に, where the second stroke failed the shape check for a short curve I consider fine.
Limits and what comes next
- One writer, the developer, writing with a finger on an iPad, one day, traced over a guide. A learner writing freehand from memory will be looser, and the thresholds were taken from Hanzi Writer rather than fitted. There are now three strictness levels, strict, standard, and loose, which scale the tolerances, the slips allowed, and how a backwards stroke counts; the next exports will show which suits freehand writing.
- Tiny strokes are judged too hard. The dots of a dakuten, as in ぐ, are a few points long, and the direction and shape checks mean little at that size. They are now judged on position alone.
- Order is reported, not enforced. The assignment is by position, so strokes drawn out of order are counted and printed but don’t change the verdict. That is deliberate for now, since a word that looks right but was written in the wrong order is a lesson rather than a failure.
- A stroke is put in the square where it starts. A stroke that wanders across the boundary into the next square is still matched against the character it started in.
Revision History
- 2026-10-10 Added the three strictness levels and the rule for dot-sized strokes; OpenAI removed from the app.