How to Train Your Dragon
If the dragon is a seven-billion-parameter Qwen. A night on the sofa with MLX: six hours, one laptop, 330,000 words of my own writing, and a model that sounds like me and knows nothing. What it learnt, what it refused to, why that is the finding, and the toolkit so you can do it yourself.

If the dragon is a seven-billion-parameter Qwen. A night on the sofa with MLX.
It had actually been a while since I had trained a model myself.
The first time I built a language model was nearly four years ago. “Built” is doing quite a lot of work there. It was really an exercise in understanding what was happening underneath all the abstraction. It took me about three months of learning to get the thing working, and the actual processing ran for several days, because the Mac I had then had no GPU to speak of.
It wasn’t particularly useful.
That wasn’t really the point.
Sometimes, if I want to understand something properly, I need to pull it apart and put it back together again. I did the same thing with operating systems, cloud platforms, distributed systems and quite a few other technologies over the years. You can read the documentation. You can watch somebody explain it. But eventually I want to know what happens when I actually turn the screws myself.
I posted about that first model on LinkedIn at the time. There were a few raised eyebrows and, I suspect, a little disbelief.
That is fine.
Four years later, what previously took months of learning and days of processing took one evening on a laptop.
That change is probably as interesting as the model itself.
And this time I had a much more practical question.
The question
AI is increasingly writing things for people.
Blogs. Journals. LinkedIn posts. Reports. Articles. Ideas turned into prose.
I don’t have a philosophical objection to that. Quite the opposite. I am slightly dyslexic, and these tools are extraordinarily useful to me. They help with ideation. They help organise thoughts. They remove some of the friction between having an idea and getting it onto a page.
But there is a problem.
Most AI writing sounds like AI writing.
If the machine is going to help me write, I don’t want it replacing me. I want it helping me sound more like me.
My phrasing. My rhythm. My slightly irritating short paragraphs. My habit of making a statement, stopping, and then coming at it again from another direction.
The machine should be the instrument, not the author.
So that became the experiment.
Could I take several hundred thousand words of things I had actually written, train a small open model on them, and give it enough of my voice that it became useful as a first-draft writing tool?
Not my knowledge.
Not my judgement.
Not a digital clone.
Just my voice.
The dragon
You cannot train a language model from scratch on one person’s writing. Pretraining needs billions of tokens; I have about a third of a million. What that much text buys you is a change of register on top of a model that already knows how to write, which is exactly the property I wanted.
So the dragon is somebody else’s dragon with my saddle on it. A seven-billion-parameter Qwen, already quantised to four bits for Apple silicon, with a small LoRA adapter trained on top: rank eight, the upper sixteen layers, eleven and a half million trainable parameters out of seven and a half billion. About 0.15 per cent of the animal.
The food was everything published on this site. Fifty-three essays. Twenty-four product pages. Five hundred glossary entries. The About page and the rest of the site’s own prose. A script reads the site repository and turns every piece into a pair: a brief in, the prose out. Write this section. Continue this section. Turn these dictated notes into this section. Define this term. Roughly 1,500 pairs, 330,000 words of target text.
Two things were kept out on purpose. The Tagalog translation, obviously. And the four-voice dossier, in which three of the voices are not mine and would have taught it the wrong accent. One essay was held back entirely and never trained on, so that at the end there would be an honest number rather than a flattering one.
Then I pressed go, at about twenty to eleven, and went to make a cup of tea.
The run
Six hours. Twelve hundred iterations, batch of two, a deliberately low learning rate, peak memory thirteen gigabytes of the thirty-six in the laptop.
Four of those six hours were lost to the Mac dozing. The trainer reported one pace and the wall clock reported a tenth of it, and the difference was the machine dropping off between wake-ups with the GPU idle. A single caffeinate pinned it awake and the rest of the run went at full speed. Write that one down; it will save you a night.

The run, drawn from its own log. Validation loss every hundred iterations; the coral ring is the checkpoint that lived.
That curve is the validation loss, measured every hundred iterations on writing the model had not trained on. It falls steeply in the first hundred, which is the model picking up the register. Then steadily. Then it flattens from 800 and stays flat. The bump at 600 was noise, gone by 700. The checkpoint at 1,000 was the low point, 1.63, and that is the one I kept.
On the essay it had never seen, the loss was higher, 2.23, which is the honest number. The validation set draws from essays the model has otherwise read; the held-out essay it has not, and the gap between those two numbers is the gap between recognising a voice and reproducing it.
Somebody else’s data centre
It is worth pausing on what the evening actually cost, because the number is the story.
The MacBook draws sixty to eighty watts flat out. Six hours is under half a kilowatt-hour, which at Bangkok rates is a couple of baht. Had I rented the equivalent instead, one H100 hour goes for two to three dollars on the marketplaces and four to six at the big clouds, and the job would have needed about one of them. The coffee I made while it ran cost more than the compute.
Now the other dragon. The one I put my saddle on.
Nobody publishes the bill for a frontier model any more. Anthropic’s launch note and system card for Fable and Mythos cover capabilities, safeguards and the knowledge cutoff and say nothing about compute, which is normal now. The last big lab to show its working was Meta, which disclosed that Llama 3.1 405B took 30.8 million GPU-hours on sixteen thousand H100s, fed fifteen trillion tokens, about fifty-four days of wall clock for the pretraining alone.
| Tokens | Compute | At rental rates | |
|---|---|---|---|
| My dragon, one evening | 1.1 million | One laptop GPU, six hours | Pence |
| Llama 3.1 405B, pretraining, as disclosed | 15 trillion | 30.8 million H100-hours | $60m to $95m |
| A 2026 frontier model, pretraining, estimated | More | Plausibly three to ten times that | $200m to $1bn |
The rental figure understates it. Post-training, the reinforcement learning and the safety work that turn a base model into something you would let near a customer, is said to be approaching the cost of pretraining for the current generation. The labs run experiments that never ship. And nobody rents at list price for two months; they own the buildings, the power contracts and the chips, and the cost is depreciation rather than an hourly rate. Dario Amodei has said publicly that current frontier runs cost of the order of a billion dollars and that the next generation will run to several, which is where the arithmetic lands too. Order of magnitude, not invoice.
So: the compute that taught my dragon its voice cost less than the electricity to boil the kettle beside it. The compute that taught it to hold a sentence together, before I ever touched it, cost somebody several hundred million dollars, took months on tens of thousands of accelerators, and was then given away under Apache 2.0.
Ten million times more tokens went into the base than into my adapter. That ratio is why a night on a laptop can teach a model a voice and could never teach it to think. The thinking was already paid for. I bought the saddle.
Four years ago that sentence would have sounded like science fiction. Now it is a line item.
What it learnt
Numbers say a style was learnt. Only reading says which. So I gave the same six lines of dictated notes to the untouched base model and to the fine-tune, with the same instructions.
Stitched the six notes into one paragraph, almost word for word. Seven sentences. Correct, complete, and lifeless.
Broke the same notes into five short paragraphs, including the two-beat construction I apparently cannot stop doing: It is not an accident. It is a habit. It also padded where it had nothing to say, and wrote one line that meant nothing at all.
That is the result, and it is the result I expected. The rhythm transferred. The short paragraphs transferred. The dryness transferred, mostly.
What did not transfer was anything resembling knowledge.
I asked it about Silo, a product with its own page in the training set. It told me, in my voice and with complete confidence, that Silo was a content management system I had built from the ground up, with Markdown and Liquid templating. Silo is a five-layer security architecture for watching AI agents. The model had read that page. It simply had not learnt to recall what a thing is from its name alone, because every training pair had put the name and the description in the brief, and the base model’s prior meaning of the word “silo” won.
Voice transfers from a corpus this size. Knowledge does not. Judgement certainly does not. And the machine will be fluent about the wrong thing without a flicker.
So I ran it again, the same afternoon, with one change to the food. Every product and every essay now had a dozen questions asked of it by name alone, “What is Silo?”, “Tell me about Saksi”, “Who is it for?”, answered from the page. And I fed it the notes I had written about the model itself, so it could describe what it is. Two thousand pairs instead of fifteen hundred, sixteen hundred iterations, five hours.

The second run, watched from the toolkit’s observability page. Blue is validation, orange is training, the grey dash is the first run for comparison, and the green lines are checkpoints. The panel on the left says what I would have said: the last point is above the low, memorising has started, keep the checkpoint nearest 1,200.
Asked about Silo by name, it now answers from the product page. Asked about Saksi, the same. Asked about AOS it invented one, and invented it out of Silo’s five layers, because no product page is called AOS and nobody had asked under that name. A data gap, not a model failure, and the aliases are in for the next run. The held-out essay came out at exactly the same loss as before, which is right: no new voice went in, only memory.
So the working rule is the one I would have written anyway. Every output is a first draft in the right register from rough notes. The facts in it are decoration until checked. An accent-removal pass strips the tells it still carries. And I read it, correct it, and sign it.
The instrument, not the author.
Where it lives
The adapter is forty-six megabytes. Fused into the base model it is under eight gigabytes, and that version loads in LM Studio and runs in Ollama on this laptop, or serves an OpenAI-compatible endpoint that my writing tools can point at as an option. Both forms are backed up to private repositories on Hugging Face.
Eight gigabytes, not four, and that number cost me an afternoon. The obvious export is to fuse the adapter straight into the four-bit base and keep it small. Do that and the voice survives but the recall does not: asked about Silo, the four-bit fuse went back to inventing a content management system, while the same adapter on the same base, unfused, answered from the product page. The arithmetic is unforgiving. A light adapter moves each weight by less than a four-bit quantisation step, so re-quantising after the fuse rounds most of the change back to where it was. The voice is spread across millions of weights and shrugs it off. The memory of one page lives in a few, and vanishes. At eight bits the step is sixteen times smaller and everything survives. Write that one down too.

The dragon in LM Studio, serving on port 1234. Four gigabytes, MLX, four-bit, and an identifier any tool on the machine can call.

The adapter’s model card on Hugging Face. The model tree on the right traces it back to Qwen.

The fused model’s files, at eight bits. Everything a Mac needs to load it.
Private, and staying private. A model that writes in one person’s voice is an impersonation kit if it sits on the public index, and I published an essay yesterday about signing what I write. The weights are mine. The method is not.
The toolkit
And because the answer turned out to be yes, with some important qualifications, I am releasing the toolkit as well.
The model is not really the interesting bit.
The interesting bit is that you can now do this yourself.

The repository. Public, MIT, and everything except the weights.
Everything except my weights is at github.com/thejustinjames/TrainYourDragon: the dataset builder that turns a folder of your writing into training pairs, the training configuration, the scripts to serve the result and export it for LM Studio and Ollama, and the documentation, including what it will and will not learn, the training notes, how to put your own model on Hugging Face, and a page on responsible use, because a tool that can imitate a voice needs one.

The README. Six commands from a folder of your writing to a first draft in your voice.

Publishing, responsible use, and the licence. The defaults are private repositories and a confirmation in words before anything goes public.
You need a Mac with Apple silicon and a reasonable amount of memory, a few hundred thousand words you actually wrote, and an evening. The rest is in the README.
Hugging Face. GitHub. Open models. Consumer hardware. Your own writing.
Take the technology and bend it toward yourself rather than allowing it to flatten everybody into the same synthetic voice.
Use AI to be more yourself, not less.
Go forth and learn.
And if you would rather somebody who has already lost a night to it trained yours, get in touch. A voice model of your own writing, on your own machine, kept to yourself.
How this was made
Dictated at ten at night, trained overnight, documented in the morning, and drafted with help from the model it describes, which is either cheating or the point. The hypothesis, the experiment and the opinions are mine. The numbers are from the training log and are reproduced, with the loss curve, in the toolkit’s documentation.