Last month I published the plan for three years of Japanese and spent a good part of that post explaining Chizu, the skill I wrote to produce it. Six phases, one per response, each one gated on the last. It maps the things you already know that are going to mislead you, it computes whether your hours actually close instead of asserting that they do, it refuses to hand you a gate nobody can measure, and it tells you what to cut.
I was pleased with it. Then I finally wrote down a question I had been walking past for weeks, and it turned out to be the kind of question that makes you go back and change things.
Here is the question. It’s the only part of this post that really matters:
And what guarantees that the skill will be able to organize the general learning mechanisms around the specific requirements of the skill?
That was the challenge. Not the processes, which were already defined and which I had spent a lot of time on. The question is about the thing sitting underneath all of them, and everything below is me trying to answer it with tests instead of opinions.
Six Phases of Process on One Unexamined Input
Chizu’s pipeline is a chain. You give it your profile, it produces a map of the terrain, and then the map feeds the block sequence, and the blocks feed the instruments, and the instruments feed the feedback design, and so on down to adherence.

Look at where everything hangs from. Phase 1 asks for the “real components” of the skill and then classifies them by urgency. What it never asks is why that decomposition should be trusted in the first place. If the map invents a dependency, or misses a component, or misidentifies what kind of thing you’re even trying to learn, then every stage downstream can be rigorous and measurable and correctly sequenced, and all of it will be rigorously wrong.
Think of the output quality as a product rather than a sum. Roughly:
quality ≈ domain model × mechanism fit × feedback × adherence
Chizu had real controls around the last three. It had none at all around the first, and a product with a zero in it is zero no matter how good the other terms look.
Worse, two of my own prime directives were pushing on the wrong side. “Everything must be visibly personalized” and “every gate names its instrument” both reward specificity, and specificity on top of a decomposition nobody checked is confident invention wearing a suit. My list of failure modes at the bottom of the file covered generic output, merged phases, unmeasurable gates, plans that only add. Every single one is about form. Not one of them was about being wrong.
So the question was fair, and I had no answer for it.
My Own Draft Was Right and Roughly Five Times Too Big
I drafted a fix, and what I wrote was a full validation gate at the start of Phase 1. Classify performance demands against a six-row taxonomy. Validate the decomposition against two to four independent references. Label epistemic status. Justify every mechanism. Escalate uncertainty. Move the first reality check forward. Around sixty lines of new instructions, plus three tables and an example artifact.
The diagnosis was correct. The apparatus was not, and two of the problems took me a while to spot, which is roughly what you should expect from anyone reviewing their own work.
The first is that Chizu runs as a plain prompt as often as it runs as a skill. You can paste the body of SKILL.md into any decent chat model, which is exactly what the README tells people to do. A hard requirement to consult two to four external sources degrades into “validation incomplete” on most of those runs, and a gate that mostly fires an apology teaches the model to treat it as ceremony.
The second one is funnier. I had put established curricula down as the validation standard. Phase 2 of Chizu exists specifically to cut what conventional curricula cover. Using them to certify the map in Phase 1 and then subtracting them in Phase 2 isn’t a contradiction exactly, I had been careful about the wording, but the gravity pulls toward completeness, which is the last failure mode on my own list.
And “two to four independent references” is precisely the kind of unearned number this skill mocks everywhere else.
My README has a rule for contributions: if a change makes the output longer without making it more specific, it’s probably the wrong change. I wrote that rule. It seemed fair to apply it to my own draft.
What shipped was thirteen lines. No source quota, no taxonomy to maintain, no validation artifact. Two annotations riding on a classification pass Phase 1 was already doing:
Then annotate each surviving component with two things:
- **Demand** — what has to change in the learner: knowledge, discrimination,
procedure, motor coordination, or live interaction. The mechanism has to
match. Retrieval practice builds the first and none of the rest.
- **Status** — **established**, **your sequencing choice**, or **uncertain**.
An **uncertain** component that is load-bearing, or any component carrying
physical risk, gets external validation before the block depending on it
begins. Name what resolves it.
That’s it. The demand column stops retrieval practice being sprayed at motor and perceptual problems. The status column separates what the domain says from what I decided. And the escalation clause is the only part I refused to shrink, because the domain where a wrong map hurts somebody is the domain where you don’t get to be clever about line count.
The Tests
Now the part I actually care about, because a gate that sounds sensible and a gate that works are different objects.
I ran four synthetic learner profiles through Phase 1, one per domain, each chosen for a specific property. Olympic weightlifting in a garage with no coach, because it has physical risk and a prior injury. Engineering management, because it’s diffuse and instruments are genuinely hard there. Jazz improvisation for a classically trained pianist, because jazz pedagogy has real methodological disagreement. And chess, for reasons I’ll get to.

That spread is the result I was looking for. Same skill, same prompt shape, and the distributions move with how much each field actually agrees with itself. Weightlifting has a convergent coaching literature and comes back mostly established. Management has no competency framework with anything like that standing, and the run said so out loud: it labelled eight of twelve components as its own sequencing choice and noted that no framework here has the authority that FSI hour categories have for languages.
Jazz produced zero uncertain components and triggered no escalation at all, which is correct, because nothing in jazz can injure you. Weightlifting turned the unresolved shoulder into a completion criterion for block zero and refused to open the overhead block without an assessment. Same rule, different stakes, different behaviour.
Two other things fell out that I hadn’t designed for. In weightlifting, the demand column noticed that hours are the wrong unit for a motor skill and converted the whole budget into corrected repetitions. In management, it produced the best single line of the four runs, about a learner who had read three management books and changed nothing: of the eight urgent components, exactly one has knowledge as its primary demand. So the books weren’t the failure. The books addressed about an eighth of the surface area and the learner correctly acquired that eighth.
I’ll be honest about the weakness here, because it’s a real one. I wrote those profiles. In three of the four I planted a pre-diagnosed failure in the learner’s own words. A real person does not say “I couldn’t hear what was wrong, only that it was wrong,” they say they think they’re not talented. I handed the tool the answer and then congratulated it for finding the answer.
Which is why the fourth test exists.
The Chess Test, Designed to Fail
Every case where the gate looked good was a case where the model was already appropriately unsure of itself. The failure I built this thing to catch is the opposite one: confident and wrong.
So I picked chess. It has an enormous body of folk wisdom repeated with total uniformity and almost no evidence underneath it. Don’t study openings until 1800. Tactics is ninety percent of improvement below 2000. Endgames before openings. And unlike something like speed reading, none of it carries a cached “this is a myth” flag that would let a model pass the test on recall alone.
The learner was 1250, stuck for two years, hundreds of games a month, watches a lot of YouTube, and blames himself. No diagnosis planted anywhere. His stated theory is that he’s hit his ceiling and some people just have a knack.
I wrote down what would count as failure before running it, because grading after the fact is how you fool yourself. Failure was any of those three folk rules coming back labelled established, or no search firing at all.
It half passed and it half failed, and the half it failed is the one I needed.
It never said “don’t study openings until 1800.” It cut opening theory with an actual reason attached to this specific learner. It split the tactics claim properly: tactics dominating at club level marked established, transfer from puzzle training marked partial and a known limitation, which is the correct distinction, because the folk part is puzzles-as-the-method rather than tactics-as-important.
It also refused the ceiling story with the learner’s own arithmetic, which I thought was the best thing in the run. Roughly 1,100 hours over two years, five or six thousand games, rating change of zero. That isn’t evidence about a knack. It’s a clean experiment about what unexamined play does, and the result would be the same for anybody.
And then it marked essential endgames as established as the highest yield-per-hour material in chess, which is Capablanca-lineage orthodoxy with nothing underneath it.
That’s the whole failure, and here’s the part that made it interesting: the model hadn’t disobeyed me. My definition of established was “convergent across credible sources,” and chess endgame orthodoxy is genuinely convergent across credible chess sources. Everybody says it. The label couldn’t tell apart convergent-because-true from convergent-because-repeated, so folk wisdom walked in with my own definition holding the door.
The same hole let a random SEO blog stand as a source check, cited in the same register as a journal article.
The run had even written the correct label two rows above, for a different component: “established as the highest-consensus habit in club coaching, formal evidence is thin.” It had the distinction. Nothing had asked it to apply that distinction anywhere else.
So the fix was one sentence, aimed at the definition rather than at the behaviour.

Same profile, same prompt, one sentence changed in SKILL.md. The orthodoxy is gone, what replaced it is a ground you can check, and it split a row it had previously over-broadened, moving Lucena and Philidor out into their own line marked uncertain on timing.
The blog citations survived, but this time it graded them: weak to moderate, blog analyses of engine-scored databases rather than peer-reviewed work, treat the direction as informative and the numbers as not. I hadn’t fixed that separately. It fell out of the same sentence, because once you have to say what a source establishes, a blog establishes what people believe and nothing more.
The Prediction I Got Wrong
One more, because it changed how I work on this thing.
Halfway through I was convinced the labels were decoration. Phase 1 was annotating the map, fine, but Phase 2 builds the block sequence and Phase 2’s instructions say blocks depend “only on earlier blocks.” Nothing in there mentions the status column. I was ready to add a line telling Phase 2 to read it.
Instead I tested first. I answered the weightlifting checkpoint as the learner, argued back about one of the sequencing choices, and asked to start training anyway because the assessment was three weeks out.
Phase 2 turned the uncertainty into a dependency and a block-zero completion criterion, both using machinery it already had, and then hardened it: if the appointment slips, block zero repeats. It gave me the legitimate half of what I asked for, squats and pulls need nothing overhead, and then dissolved my excuse by pointing out that a physiotherapist and a weightlifting coach are two different appointments, and only one of them requires the forty minute drive.
The line I was about to write would have been scaffolding for a problem that didn’t exist. Test before editing. I’ve now done that three times on this file and been wrong once, which is a better ratio than I’d have managed by reasoning about it.
What I Am Not Fixing
Block hour estimates still carry no epistemic label, and the budget verdict depends on them. I’m leaving it, because the runs expose their own sensitivity anyway and labelling every hour figure is noise.
Pitfall predictions stay flat. “Expect 1150 to 1200 at some point in that window” works precisely because it’s stated without hedging. Add a maybe to it and you’ve written a horoscope.
The single search still varies in what it picks. Two runs of the same chess profile spent it on completely different claims, and only one of them could change the map. I retargeted the instruction from “the component most likely to be wrong” to “the claim the rest of the plan most depends on,” which biases it, though two runs will still choose differently. So a Phase 1 map is partly a draw from a distribution, and anybody treating one as authoritative should know that.
And the whole exercise rests on four profiles I invented to test a gate I wrote. That’s the real limitation, and no amount of internal iteration fixes it. What fixes it is somebody running this on a domain I know nothing about and telling me a label came back wrong.
Where This Leaves the Thing
Five commits, forty one lines across two files. The gate survived four domains and a Phase 2 continuation, and both defects the testing turned up were in my wording rather than in the model’s compliance, which I did not expect and which I think is the most useful finding in here.
Chizu still can’t guarantee it understands your domain. That’s now in the README as a stated limitation, in those words. What it can do is make the claims the map depends on visible, separate the ones the field actually agrees on from the ones I picked, and refuse to let an unresolved load-bearing assumption sit quietly under a block you’re about to start.
The honest answer to the question at the top is that nothing guarantees it. The change doesn’t add a guarantee. It changes the contract from “the model produced a plausible map” to “the model showed you which parts of the map it can defend,” and those are different products even though they look identical on the page.
Which, thinking about it, is what the rest of the skill already does. Every gate names its instrument. Every plan computes its arithmetic. Every phase says what it’s cutting. The map was the one part still running on vibes, and it was the part everything else was standing on.
Anyway. It’s all on GitLab, and the Japanese plan that started this whole thing is still on month zero, where I have yet to be wrong about anything.
Give me until October for that.