Back to Interactive Learning
Advanced Topics

Large Language Models Deep Dive

Explore how LLMs are trained from scratch - from collecting massive datasets to fine-tuning with human feedback. Follow the interactive step-by-step journey.

Interactive Learning

How Large Language Models Are Trained

Training an LLM is a complex, multi-step process that takes months and millions of dollars. Click through each step to understand how it works.

STEP 1 OF 6

Data Collection

Gather massive amounts of text data from books, websites, articles, and more.

Modern LLMs are trained on trillions of words from diverse sources including books, scientific papers, code repositories, websites, and conversations.

Training Data Sources:

Clinical Notes & EHR Records

500 million records

≈ 800B tokens

Medical Literature (PubMed)

35 million articles

≈ 400B tokens

Radiology & Pathology Reports

200 million reports

≈ 300B tokens

Medical Textbooks & Guidelines

80,000 texts

≈ 150B tokens

1 of 6

Key Takeaways

  • • Training LLMs requires massive computational resources (10,000+ GPUs, months of training)
  • • Models learn by predicting the next word billions of times, gradually improving
  • • Fine-tuning with human feedback makes models more helpful and safer
  • • The quality of training data directly impacts the model's capabilities
  • • Even after training, models continue to be updated and improved