Explore how LLMs are trained from scratch - from collecting massive datasets to fine-tuning with human feedback. Follow the interactive step-by-step journey.
Training an LLM is a complex, multi-step process that takes months and millions of dollars. Click through each step to understand how it works.
Gather massive amounts of text data from books, websites, articles, and more.
Modern LLMs are trained on trillions of words from diverse sources including books, scientific papers, code repositories, websites, and conversations.
500 million records
≈ 800B tokens
35 million articles
≈ 400B tokens
200 million reports
≈ 300B tokens
80,000 texts
≈ 150B tokens