Training Not So Large Language Models on My PC
In my post The Big Learning Set for Big Word Helpers I explained how I planned to generate a training dataset using the 1000 most common English words from the original Up Goer Five xkcd comic, and how I had started doing so with a custom agent provided with tools to check and refine its own generations. Here I’ll go over the continuation of those efforts, up to a live demo of a language model (“large” at this point may be a misnomer) trained on this dataset!
Expanding the dataset#
I started by training some simple models on what I had generated manually or via the agent described in my earlier blog post, which amounted to about ~70k words. I adopted a very classic model architecture, straight out of the original GPTs, with two different size presets, gpt-small (a bit less than 1M parameters) and gpt-tiny (about 170k parameters):

The results were not great, though there was clearly some learning going on. I ended up leaving the project aside for a while, then only recently picking it back up. This time, I adopted a different approach. First, I added the generating agent a set of tools to consult a dictionary of standard circumlocutions - expressions using the allowed vocabulary to describe things that aren’t in it. For example “law” is “the things everyone must follow”, or “garden” is “the growing ground outside a home”. Second, and most important, I used the existing data to finetune a local model to speak only in the simpler language to begin with. I did this in two steps:
- first, I went over the existing generation with LFM2.5-2.6B to generate a question they could plausibly be an answer to;
- then I finetuned a Llama3.2-1B instruct model on the resulting instructions with Unsloth.
The result is a model I dubbed “tinyfacts-llama”. Here is an example generation from it:
User: Explain the story of the movie Casablanca. Do not use any proper nouns of places or people.
Model: There was a man who lived a life that was full of many different kinds of love. He had a friend who was like a brother. They did many things together. They also had a love for a girl. She was very sweet. The man loved her but he did not know how to say it. He felt a deep kind of pain inside him. One day the man was in trouble. A man came into his home and asked for money. The man gave it to him. The man did this because he wanted to help his friend. But he was also afraid. He knew the man was going to do something bad with the money. The man met a woman. She was the girl he loved. He wanted to tell her how he felt, but he was afraid. He was afraid of losing her. He wanted to keep her safe. The woman told the man to go to a place where he could get money without having to deal with the bad man. The woman said he should go there because he would be safe. The man went to the place and found out that the bad man had taken the money. He also found out that the woman he loved was in trouble because of what he did. The man decided to go back and try to help his friend. He went to the woman and said he was coming back. He told her he wanted to see her. The woman said they should meet at the old house on the road. The man went there and found her. They were both scared, but they found a way to be together. The man said he would do anything for the woman. He would give his life if he had to. The woman was glad. She knew the man was a true friend. They were happy to be together again. But the bad man did not give up. He kept trying to find the man. He sent his friends to catch him. The man was in a lot of trouble. He had to find a way to get away from the bad man and his friends. The man came up with a plan. He told his friend to help him. Together they were able to escape. The man and his friend went to a place where they could hide. The bad man and his friends were there too. The bad man was angry. He wanted to kill the man. But the man was ready. He had planned his death. The bad man was killed, and the man was safe. The woman and the man were finally able to be together. They were happy. They knew they had been through a lot, but they had come out stronger. They loved each other, and that was the most important thing of all.
The language is coherent enough and resembles the existing generations by high end models, though the specific recollection of a beloved cinema classic leaves something to be desired (some plot beats are correct, but the overall picture isn’t). Using this model, llama.cpp and about ~8 hours on my desktop PC to ask it to write individual explanations for each of the 10,000 most common English words I increased the size of the dataset from ~70k words to almost ~3M. While I’m currently keeping the model private, the final dataset is available publicly on Huggingface.
Tokenizer#
The 1000 words plus their various forms (verb pasts, noun plurals, etc) totaled to almost 3k words. In order to keep the vocabulary smaller and exploit regularities in patterns I executed my original idea of grammar-based tokens. This means in practice supporting the following special tokens:
_SUPPORTED_INFLECTIONS = {
'NNS', # Noun, plural
'VBZ', # Verb, 3rd person singular present
'VBD', # Verb, past tense
'VBG', # Verb, gerund or present participle
'VBN', # Verb, past participle
'JJR', # Adjective, comparative
'JJS', # Adjective, superlative
'RBR', # Adverb, comparative
'RBS' # Adverb, superlative
}
which can modify certain specific words and thus produce the declined forms. For example [VBD]sleep becomes slept. This way, the total effective vocabulary can fall even below 1000 words (some of the original 1000 words were already themselves derived forms, and some ended up never appearing in the training set).
Training and results#
I trained both the gpt-small and the gpt-tiny architectures for 100,000 steps, each corresponding to a batch of 256 sequences of 128 tokens each. I kept 5% of the dataset aside for validation and the rest went into training, for a total of almost 1000 epochs. The result differed. gpt-tiny seemed to be small enough to never saturate, even with that many repetitions:

gpt-small instead saturated around 30k steps, at which point training loss kept decreasing but validation loss started increasing, a sign of overfitting:

Here are two example generations. Using the start “a city is”, gpt-small completes:
a city is a group of people. they work together to make a big group more money.
they help each other. they help each other. they are also the ones who help
each other. they are the ones who help each other. sometimes, a group of
people who are friends is
Meanwhile, gpt-tiny:
a city is a place with many people. they go to the city and to see many places.
a city is a place where people go to see the city. a city is a place where people
go to go to a far place. they go to a far place to go to the city.
Not too bad for something so small! Ultimately:
- they both use grammatically correct, if simple, English;
- they seem to have the right cluster of ideas and even a sense of what the words mean (though some of it is memorization, it’s not verbatim;
tinyfacts-llama’s generation on the word city begins withA city is a place where many people live and work.); - they still fall prey to repetition, redundancy, and quickly lose coherence.
This is not as good as the Tinystories models, but gpt-tiny is also 1 OOM smaller than the smallest model of that series, and trained on about 3 OOM less text.
Online playground#
Finally, you can experiment with the models too! I exported them to ONNX and used a web runtime to have them execute entirely in the browser; they’re just a few MB large, after all. Here is the live playground.
Future steps#
Future steps definitely involve two things:
- increase the size of the dataset;
gpt-smallclearly has more potential to learn if it’s fed an appropriate diet of more tokens; - experiment with architectural changes. I’ve tried already Mamba and other architectures in the past and mostly desisted because they were a lot more inefficient to train on my hardware, but perhaps implementation can be improved. And some other changes, like e.g. the use of RoPE, would ditch some parameters and make the models even smaller