Pretraining LLMs: data, objectives and compute budget
When you talk to Claude or GPT-4, you are interacting with a model that went through two distinct phases. The first — and by far the most expensive — is pretraining: exposing the model to trillions of tokens of text and teaching it one simple task. The second phase (alignment, covered in the next lesson) turns that raw capability into a helpful assistant.