Teach With This Book¶
Everything a course needs is here and free: the text, a lecture deck for every chapter, an assignment for every chapter, and notebooks that run in a browser. All of it was written for business students who have not programmed before.
What Comes With Every Chapter¶
- A lecture deck. Timed for one 75-minute class, with speaker notes. Every number on a slide is one the chapter's own script printed. See all 17 decks.
- An assignment. About 15 minutes. The student runs a script, changes one number, and explains what changed. There is no new code to write.
- A notebook. The chapter opens in Google Colab with the code ready to run, so no student is stopped by an installation.
Every chapter also opens with the terms it introduces and ends with a summary and a short set of Check Your Understanding questions. The glossary collects the terms in one table.
A Course, Class by Class¶
The decks are timed for one chapter per class, so the book is 17 classes. A course that meets twice a week covers it in nine weeks, which leaves room in a term for a project and for exams.
Table T.1: One class per chapter, with its assignment, deck and notebook.
The first class is a setup lab, not a lecture. Students who use the browser notebooks can skip the installation and join at Chapter 2.
What a Student Can Do After Each Class¶
Each chapter's goals are written here as things a student can do, next to the assignment that shows it. The rows can be pasted into a syllabus.
Table T.2: What a student can do after each class, and the assignment that shows it.
| Class | After this class a student can | Shown by |
|---|---|---|
| 1. Environment Setup | Set up Python, the project code and an isolated workspace, and confirm the setup with the book's check script. Explain why a team checks its setup before work starts. |
Assignment 1 |
| 2. What Is an LLM? | State the one rule a language model follows, which is to predict the next piece of text. Describe how a model builds a long answer one prediction at a time. Explain how plain text supplies its own practice questions and answers. |
Assignment 2 |
| 3. Tensors and PyTorch | Read a tensor's shape and say what each dimension holds. Predict the shape a matrix multiplication produces. Turn raw scores into probabilities with softmax. |
Assignment 3 |
| 4. Tokenization | Explain why text must become numbers before a model can use it. Encode text into token IDs and decode it back. Explain why a model treats two spellings of one word as different text. |
Assignment 4 |
| 5. Embeddings | Explain that an embedding is a list of numbers looked up by token ID. Show that the size of a token ID carries no meaning. Explain why position information is added. |
Assignment 5 |
| 6. Self-Attention | Describe attention as a weighted average over earlier tokens. Name the roles of the query, the key and the value. Explain what the causal mask hides and why. |
Assignment 6 |
| 7. Multi-Head Attention | Explain why a model runs several attention heads at once. Compare what different heads attend to in a trained model. Explain why the heads' outputs are joined end to end. |
Assignment 7 |
| 8. Feed-Forward and Norms | Describe what the feed-forward network does to each token after attention. Compare GELU with ReLU on positive and negative numbers. Explain why layer normalization keeps the numbers stable. |
Assignment 8 |
| 9. The Transformer Block | Name the parts of a Transformer block in order. Explain what a residual connection keeps and what it adds. Explain why blocks can be stacked. |
Assignment 9 |
| 10. The Full GPT Architecture | Trace an input from token IDs to scores through the full model. Read the model's parameter count and say where most of the parameters sit. Explain what logits are and why they are not yet probabilities. |
Assignment 10 |
| 11. Causal Language Modeling | Build input and target pairs by shifting a sequence one position. Read cross-entropy loss as surprise at the correct answer. Describe generation as a loop of predicting and appending. |
Assignment 11 |
| 12. Dataset and DataLoader | Explain how a sliding window turns one text into many training examples. Explain what a batch is and why batches speed up training. |
Assignment 12 |
| 13. The Training Loop | List the four steps of the training loop and say what each does. Predict what a learning rate that is too small or too large does to training. Read a loss value as the probability the model gave the correct character. |
Assignment 13 |
| 14. Checkpointing | Save a trained model and load it back. Say what a checkpoint file must hold besides the weights. Explain why a model is switched to evaluation mode before use. |
Assignment 14 |
| 15. Greedy and Sampling | Generate text with greedy decoding and with sampling. Explain why greedy decoding repeats itself and sampling varies. Choose a decoding rule for a task that must give the same answer every time. |
Assignment 15 |
| 16. Temperature and Top-k | Predict how lowering or raising the temperature changes the output. Explain what top-k removes from the choice. Recommend settings for an assistant that writes to customers, with reasons. |
Assignment 16 |
| 17. Putting It All Together | Run the full pipeline from text to generated output. Explain why the small model continues text and does not answer questions. Say what separates this model from a production assistant. |
Assignment 17 |
How Students Follow Along¶
- In class. Every deck has a slide where the instructor runs one of the book's scripts in front of the class. The numbers on the slide are what the script prints.
- In a browser. The button at the top of every chapter opens that chapter's notebook. The whole script sits in one cell, ready to edit and run.
- On a laptop. Chapter 1 installs Python and the project step by step. Training needs no graphics card.
The Assignments¶
Every assignment has the same four steps, so students learn the routine once.
- Run it. Run the chapter's script as it is, and copy down what it prints.
- Change one number. Change one value in the script (one number, or one word where the script works on text), and run it again.
- Compare. Say what changed and what did not.
- Explain it. Write two or three sentences a manager could follow.
Every hand-in therefore has the same three parts: the two runs, the comparison and the explanation. One rubric serves them all. Chapter 1 is the one exception. It is the setup chapter, so its assignment asks the student to run the setup check and read its output, with nothing to change.
Table T.3: One rubric for every assignment. The points add up to 10.
| Part of the hand-in | Full credit | Half credit | No credit |
|---|---|---|---|
| The two runs (3 points) | Both outputs are there, copied as the script printed them. The second run shows the one change the assignment asked for. | One output is missing or retyped, or more than one thing was changed. | No output from the student's own run. |
| The comparison (3 points) | Every question in step 3 is answered, and each answer points to a line of the output. | Some questions are answered, or the answers do not use the output. | No comparison. |
| The explanation (4 points) | Two or three sentences a manager could follow that say why the output changed. | The sentences say what changed and leave out why. | No explanation, or one copied from the chapter. |
For Chapter 1 the second run is the output printed in the chapter, and the comparison is between the student's computer and the book's.
Slides That Are Easy to Follow¶
Every slide has a kind, marked by a color down its left edge and a chip in its corner: an idea, code to read, a run, the student's turn, a business example, a common mistake, the summary, the assignment. No slide carries more than 85 words of prose. The slides page shows the colors.
The License, in Plain Words¶
The book text and the lecture decks are licensed CC BY-NC-ND 4.0. You may assign them, print them and share them unchanged with your students, with credit and for noncommercial teaching. You may not distribute a changed edition or a translation without permission. The code and the notebooks are under the MIT License, so you and your students may change and reuse them freely. The exercises in the companion repository are under CC BY-NC 4.0, so you may adapt them for your own course.
How to Cite¶
Luu, T. (2025). Teach your computer to write: Build and train an LLM from zero (October 2026 ed.). https://jackluu.io/book/
In BibTeX:
@book{luu2025teach,
author = {Luu, Truong},
title = {Teach Your Computer to Write: Build and Train an {LLM} from Zero},
year = {2025},
edition = {October 2026},
url = {https://jackluu.io/book/}
}
If You Adopt the Book¶
Please write to hello@jackluu.io with the course name and the term. Knowing who teaches from the book decides what is improved next. Corrections are welcome through the same address or as an issue.