6. The First Model That Actually Learns
Summary
Least squares on four points
A model of this kind is a weighted sum of features plus a rule for choosing the weights. Four invented points (hours studied against quiz score) are fitted by hand. Each error is squared so misses can't cancel and big misses cost more. The line that minimises the sum of squared errors comes out as prediction = 1.4x + 0.5, with squared errors totalling two tenths. scikit-learn's LinearRegression returns the same coef_ and intercept_. It needs x as a four-by-one column, because the library always expects a two-dimensional feature table. Squaring is flagged as a choice that suits symmetric noise with rare large misses. A later chapter gives that assumption a precise form.
From a line to class probabilities
A line can't name a digit. An output of 4.6 is no class, and nothing keeps predictions in range. Logistic regression keeps the weighted sum but changes what is done with it. A sigmoid squashes the sum into a yes-or-no probability. Softmax turns ten per-digit scores into ten probabilities that add to one. Scoring switches to log loss. It charges about a tenth for giving the right digit 0.9, about 0.7 for 0.5, and about 4.6 for 0.01. Log loss has no closed-form solution, so the solver improves the weights step by step.
Ninety-two percent on MNIST
LogisticRegression goes inside a pipeline after StandardScaler. That way the scaler learns averages from training images only. Scaling evens out the loss surface so the lbfgs solver converges faster. The default max_iter of 100 triggers a ConvergenceWarning on MNIST, so the chapter uses 1,000. The deprecated multi_class parameter should be left out. The test score lands around 92–93 percent, against the 11.35 percent baseline.
Per-digit accuracy is uneven, and listeners are told to read their own numbers. Reshaping the coefficient row for zero to 28×28 shows a rough ring of positive weight with a negative centre. That template shows both what the model learned and its limit with shifted strokes. The exercise compares scaling choices and iteration limits on a validation split, never on the test set.
Python 3.15 delayed a week
Blocking bugs in the lazy-import feature (PEP 810) pushed the final release from October 1 to October 9. Release manager Hugo van Kemenade issued a third release candidate on October 2. pandas 3.0.6 and scikit-learn 1.9.1 already ship prebuilt packages for 3.15, in standard and free-threaded builds. Once 3.15 final is out, the suggested check is to pin a copy of the project to it with uv, confirm the library versions, and rerun the notebook. Polars 2.0.0 makes streaming the default for lazy queries. Its joins and group-bys no longer preserve row order unless maintain_order is set, which can silently misalign positional splits.
The last chapter ended with a number every model has to beat. A model that guesses "one" for every MNIST test image gets about 11.35 percent right. The stratified and uniform versions of that guess are still an open exercise, and both should land near ten percent. This chapter fits the first model that learns something from the pixels. The question is what that model is and how it picks its numbers.
The answer is short. A model like this is a weighted sum of features plus a rule for choosing the weights. Linear regression and logistic regression both have that shape. They differ in two ways: what the sum is used to predict, and how the model is charged for being wrong. Linear regression comes first because it is small enough to do by hand.
A line chosen by its errors
Take four invented points. Think of x as hours studied and y as a quiz score. The points are x equals one with y equals two, x equals two with y equals three, x equals three with y equals five, and x equals four with y equals six.
The model predicts with one formula: prediction equals w times x, plus b. Here x is the input. The weight, w, says how much the prediction rises when x goes up by one. The bias, b, is where the line sits when x is zero. Choosing a model means choosing w and b.
To choose them we need a way to score a line. For each point, the error is the true y minus the predicted y. Some errors come out positive and some negative, so adding them as they are would let them cancel. Squaring each error makes every one positive and charges big misses much more than small ones. A miss of two costs four, while a miss of one costs one. Add the squared errors over all the points and you get one number for the whole line, the sum of squared errors. The best line is the one that makes that sum as small as it can be. That is what least squares means.
The smallest sum has a closed-form answer. For one input, w equals the sum of the products of how far each x is from the average x and how far each y is from the average y, divided by the sum of the squared distances of x from its average. Then b is whatever makes the line pass through the point made of the two averages.
Now work it through. The average x is two and a half. The average y is four. The x distances from average are minus one and a half, minus a half, plus a half and plus one and a half. The y distances are minus two, minus one, plus one and plus two. Multiply them in pairs: three, a half, a half and three, for a total of seven. The squared x distances are two and a quarter, a quarter, a quarter and two and a quarter, for a total of five. So w is seven divided by five, which is one point four. Then b is four minus one point four times two and a half. That is four minus three and a half, which is one half.
The line is: prediction equals one point four times x, plus one half. Check it against the data. At x equals one it predicts one point nine against a true two. At x equals two it predicts three point three against three. At three it predicts four point seven against five. At four it predicts six point one against six. The errors are a tenth, minus three tenths, three tenths and minus a tenth. Their squares add to two tenths. No other straight line gets a smaller sum on these four points.
scikit-learn gives the same answer. In the project's notebook, import LinearRegression from sklearn dot linear_model. Make x a column of shape four by one, because scikit-learn always expects a two-dimensional feature table, even when there is one feature. Then call LinearRegression, then fit with x and y. The fitted model's coef_ attribute holds the weight, and it shows one point four. Its intercept_ attribute holds the bias, and it shows zero point five. The predict method gives back the same four predictions worked out above. The library is doing the same minimization at any size.
Squaring was a choice, and it will come back. Squared error treats a miss upward and a miss downward the same, and it punishes large misses hard. That fits best when the noise around the true line is symmetric and large misses are rare. Later in the course that assumption gets a precise form, a bell-shaped noise model. Under that model, least squares turns out to be the most likely line. For now, squaring is a reasonable way to keep errors from cancelling.
Why a line cannot name a digit
Try this machinery on MNIST. The features are 784 pixels, so the model needs 784 weights and one bias, and it produces one number. But the target is a class. A label of seven is not "more" than a label of three. If a line predicted four point six, it would be naming something between a four and a five, and that is not a digit. A line will also predict minus twelve or ninety if the pixels push it there. Nothing keeps the output in the range a class answer needs.
The fix keeps the weighted sum and changes what is done with it. Start with a yes-or-no question: is this image a seven? Compute the weighted sum of the pixels plus a bias. That sum can be any number, so pass it through the sigmoid function. The sigmoid takes any number and returns a value between zero and one. A large positive sum comes out close to one. A large negative sum comes out close to zero. A sum of exactly zero comes out at one half. Read that value as the model's probability that the answer is yes.
Ten digits need ten answers at once, so give each digit its own weight vector. That is ten rows of 784 weights, plus ten biases. Each image now gets ten weighted sums, one score per digit. Softmax turns those ten scores into ten probabilities. It raises the constant e to the power of each score, which makes every score positive and makes big scores much bigger. Then it divides each result by their total, so the ten probabilities add to one. The prediction is the digit with the highest probability. With two classes, softmax reduces to the sigmoid.
The scoring rule also has to change. Squared error on probabilities works poorly, because it barely separates a hesitant mistake from a confident one. Logistic regression uses log loss. For each image, look at the probability the model gave to the correct digit and take the negative of its logarithm. If the model gave the right digit a probability of nine tenths, the cost is about a tenth. If it gave it one half, the cost is about seven tenths. If it gave it one in a hundred, which means it was confidently wrong, the cost is about four point six. The cost has no upper limit as that probability falls toward zero. So log loss charges little for being right, a moderate amount for hedging, and a great deal for being sure of the wrong answer. Training searches for the ten weight vectors that make the average log loss over the training images as small as possible. The information-theory phase will give this loss its formal name and meaning, cross-entropy.
There is one practical difference from the line. Least squares had a formula we could solve by hand. Log loss has none, so the library starts with some weights and improves them step by step. That makes the starting scale of the data matter, as the next part shows.
Fitting logistic regression on sixty thousand digits
The data route is the one already in place: fetch_openml with mnist_784, version one, with return_X_y and as_frame set as before. The first 60,000 rows are for training and the last 10,000 for testing. The test rows stay untouched until the end.
The model is scikit-learn's LogisticRegression, but it does not run alone. It goes inside a Pipeline with a scaler in front. The call is make_pipeline from sklearn dot pipeline, given StandardScaler first and then LogisticRegression with max_iter set to one thousand and random_state set to forty-two. StandardScaler shifts each pixel column so its average over the training images is zero and its spread is one.
Scaling matters because of how the weights are found. The default solver in current scikit-learn is called lbfgs. It improves the weights in repeated steps, each step guided by the slope of the loss. Raw pixels run from zero to 255. Inputs that large make the loss very steep in some directions and nearly flat in others. A solver that walks in steps then makes slow, zig-zagging progress. When every pixel has a similar scale, the loss surface is more evenly shaped, and the solver gets to the bottom in far fewer steps. Dividing every pixel by 255, to put it between zero and one, helps for the same reason. That is a reasonable alternative to try.
The scaler goes inside the Pipeline for a reason tied to the last chapter. The scaler learns averages and spreads. It must learn them from training images only. When you call fit on the pipeline with the training arrays, the scaler is fitted on those images and the model is trained on the scaled result. When you later call score with the test arrays, the pipeline applies the training averages to the test images. It never recomputes them. The test set does not leak into the fit.
Some pixels are a special case. Border and corner pixels are zero in nearly every training image, so their spread is zero. StandardScaler handles this by leaving those columns at zero rather than dividing by zero. They simply contribute nothing.
Two settings deserve a word. The first is max_iter, the most steps the solver may take. The default is one hundred. On MNIST that is not enough, and scikit-learn shows a ConvergenceWarning that says lbfgs failed to converge. The warning means the solver stopped before reaching the minimum, so the weights you got are unfinished. Raising max_iter to five hundred or more, here one thousand, gives it room to finish. The second is the old multi_class parameter, which many tutorials still set. It is deprecated in current scikit-learn. Leave it out. With a target of ten classes, the library already fits the softmax version described above. If you want ten separate yes-or-no models instead, wrap the estimator in OneVsRestClassifier.
On an ordinary laptop processor, the fit usually takes somewhere between under a minute and about three minutes. The model learns ten weight vectors of 784 numbers each, plus ten biases, from sixty thousand images.
Then comes the one score on the test set. Call score on the pipeline with the test images and test labels. A model of this kind lands at about ninety-two to ninety-three percent accuracy on those ten thousand images. Put that beside the floor. The baseline got about one image in nine right. This model gets about nine in ten right. Every percentage point between 11.35 and 92 comes from the weighted sum reading pixels.
Overall accuracy hides how each digit fares, so compute per-class correctness the same way as for the baseline. Use predict on the test images to get the predicted labels. Then, for each digit from zero to nine, take the test images whose true label is that digit and find what fraction were predicted correctly. The baseline scored one hundred percent on ones and zero on everything else. This model gives every digit a real score, and the scores are not equal. Digits with simple, consistent shapes tend to score highest, while digits that share strokes with other digits, like a five that resembles a three, tend to score lower. Read your own ten numbers rather than trusting that sketch. They are the start of an error analysis, and the confusion matrix later shows which digits get mistaken for which.
The last step is to look at what was learned. The fitted classifier sits inside the pipeline under the name logisticregression. Its coef_ attribute is a table of ten rows and 784 columns, one row per digit. Take one row, say the row for zero, reshape it to 28 by 28, and show it with matplotlib using a color map where positive is one color and negative another. A rough ring shows up. Pixels where a zero's stroke usually falls carry positive weight and push the score for zero up. The center, where a zero has a hole and many other digits have ink, carries negative weight. The corners are blank, because those pixels never change. This image is the whole model for that digit: a template it compares against every image, pixel by pixel, then adds up. It also shows the limit of the model. One fixed template per class cannot follow a stroke that has shifted a few pixels to one side. That limit is why later models move beyond it.
The exercise works on training data only. Split off the last ten thousand of the sixty thousand training images as a validation set, and fit on the first fifty thousand. Then change one thing at a time. Swap StandardScaler for dividing by 255. Fit without any scaling. Try max_iter at one hundred, three hundred and one thousand. For each run, write down whether the convergence warning appeared, how long the fit took and the validation accuracy. Predict the results before running them. Expect the unscaled, low-iteration run to warn and to score lower. Do not touch the test set while doing this. It has already been scored once. Checking it again while comparing settings would make it a tuning set, and then it would no longer measure new handwriting.
Python 3.15 waits one more week
The last news item from the series tracked the second release candidate of Python 3.15. The final release did not ship on October first as planned. Last-minute blocking bugs in the new lazy-import feature, Python Enhancement Proposal 810, led release manager Hugo van Kemenade to delay it by one week. He put out a third release candidate on October second and moved the final release to October ninth.
The libraries are already prepared. pandas 3.0.6, released in mid-September, has prebuilt packages for Python 3.15, in both the standard build and the free-threaded build. scikit-learn 1.9.1, from September tenth, has both as well. Prebuilt packages matter because without them your machine has to compile the library from source, which is slow and often fails.
The practical step is the one from the earlier item, and it should be done in a copy of the project, not the working one. Once 3.15 final is out, have uv pin the copy to Python 3.15 and run uv sync. Use uv tree to confirm pandas and scikit-learn resolved to those versions. Then restart the kernel and run all of this chapter's notebook. If the accuracy and the per-class scores match the 3.14 run, the upgrade is safe for this project. If they move, the change is real and needs explaining before you adopt it.
One more release is worth knowing about if you use Polars, the dataframe library often used instead of pandas. Polars 2.0.0 shipped on October sixth. It makes its streaming engine the default for lazy queries, which saves memory. Joins and group-bys no longer promise to keep row order unless you ask for it with maintain_order set to true. Code that relied on order after a join, such as positional splits, can now quietly misalign rows. That is the same kind of leak the train and test split was built to prevent.
