Training an LLM on Nothing But 1800s London Text
Most language models you interact with carry the fingerprints of the modern internet. They know about smartphones, memes, and current events, and when you ask them to sound like a Victorian, they're really just doing an impression of one. What if a model didn't pretend to be from the past, but actually was shaped by it? TimeCapsule LLM is an attempt to answer that question by training a language model from scratch exclusively on data from a specific place and time period.
What It Does
TimeCapsule LLM is a language model trained from scratch on data drawn only from certain places and time periods, with the goal of reducing modern bias and emulating the voice, vocabulary, and worldview of that era. The current focus is London between 1800 and 1875. Because the training data is restricted to that window, the model's outputs reflect the language and thinking of the period rather than contemporary text.
The project has gone through several iterations. The v0 and v0.5 versions are built on Andrej Karpathy's nanoGPT, using his core training scripts and model architecture. The v1 model is built on Microsoft's Phi 1.5, and v2 is built on LlamaForCausalLM. Models and datasets are hosted on Hugging Face under the TimeCapsuleLLM 1800–1875 London collection.
The project was initiated and developed independently, and it's now conducted under academic supervision with an affiliated research collaboration at Muhlenberg College and Georgia State University. If you use the dataset or model in academic work, there's a BibTeX citation provided for the Historic London English (1800–1875) dataset.
Why It's Cool
-
The constraint is the whole point. Plenty of projects fine-tune a modern model on historical text, but the modern bias never fully goes away. By training from scratch on a narrow slice of history, TimeCapsule LLM sidesteps that contamination entirely. It's a clean experiment in what a model becomes when its entire world is 1800s London.
-
The early results are genuinely strange in the best way. The README shares a v0 example where the prompt "Who art Henry?" produced the response "I know that man, I have did not a black, the storm." That's not a polished chatbot answer—it's the kind of output you'd expect from a model that has only ever read period text. The grammar is off, but the voice is unmistakably of the era. That's the interesting part.
-
It's built on well-known foundations. Rather than inventing an architecture from nothing, the project leans on nanoGPT, Phi 1.5, and LlamaForCausalLM across its versions. That makes the work easier to follow if you already know those codebases, and it keeps the focus on the data and the era rather than the plumbing.
-
There's a community around this niche. The README points to a "Vintage LLM" Discord for people interested in historical language models, time-specific datasets, and related projects like Violet-1.4B and Mr. Chatterbox. If this is your kind of thing, you're not alone in it.
-
The academic framing adds weight. Being conducted under supervision with a college and university collaboration suggests this isn't just a weekend experiment—there's a research angle here around historical language modeling and dataset construction.
How to Try It
The models and datasets live on Hugging Face, so that's your starting point.
-
Head to the Hugging Face collection for TimeCapsuleLLM 1800–1875 London: huggingface.co/collections/haykgrigorian/timecapsulellm-1800-1875-london
-
Grab the model version you want to experiment with. Remember the version differences: v0 and v0.5 use nanoGPT, v1 uses Phi 1.5, and v2 uses LlamaForCausalLM.
-
If you're using the dataset in academic work, cite it properly:
@misc{london_llm_1800,
author = {Grigorian, Hayk and Yaghoobian, Hamed},
title = {Historic London English (1800–1875)},
year = {2025},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/datasets/postgrammar/london-llm-1800}}
}
-
The full source and training scripts are on GitHub: github.com/haykgrigo3/timecapsulellm
-
If you want to talk with others building in this space, the Discord link is in the README.
Final Thoughts
TimeCapsule LLM is a narrow project, and that's its strength. It's not trying to be a general-purpose assistant—it's trying to be a faithful artifact of a specific time and place, and the from-scratch training approach is what makes that credible. The outputs shown so far are rough around the edges, which is exactly what you'd expect from a model with such a constrained diet. If you're into historical language modeling, dataset design, or just want to see what happens when you strip modern bias out of an LLM entirely, this is worth following. The project is still evolving through its versions, and the research collaboration suggests there's more to come.
Follow @githubprojects for more developer tools and open source projects.