5. The Number Every Model Must Beat
Summary
Reading the arrays before any model
X.shape prints (70000, 784): seventy thousand images, one per row, and one pixel position per column. y.shape prints (70000,), a flat list with one label per row of X. Checking dtype shows pixel values from 0 to 255 and labels stored as text strings, the character "5" rather than the number five. Slicing at row sixty thousand gives a 60,000-image training set and a 10,000-image test set with no overlap. Reshaping one row to 28 by 28 shows the stroke of a handwritten five.
pandas is used only for inspection. value_counts shows the classes are close to balanced. Ones are the most common class at about eleven percent. A small per-pixel table shows that corner pixels are almost always zero, so some features carry almost no information. The model still trains on the raw NumPy array, not on this table.
Keeping the test set shut
The test slice is set aside before any modeling choice and scored once, at the end. Checking it repeatedly while tuning means the final number measures fit to those ten thousand images, not to new handwriting. Slicing by position works for MNIST because the dataset was built with this split in mind. Most datasets have no such guarantee. The general tool is a shuffled, stratified split with a fixed random seed, as described in the train_test_split reference. A random split can still leak information through repeat writers, duplicate rows or repeated customers. The fix is to split by the unit that has to stay separate.
A model that ignores the pixels
DummyClassifier with most_frequent looks only at the labels and predicts "1" for every image. Its test accuracy is 1,135 correct out of 10,000, about 11.35 percent. Broken down by digit, it is perfect on ones and gets every other digit wrong, a pattern the single accuracy figure hides. An exercise compares this with the stratified and uniform strategies. Both should land near ten percent and spread their errors evenly across the digits. most_frequent wins on accuracy and is also the most lopsided. Linear and logistic regression come next, and their scores are to be read against this floor.
Python 3.15 and a pinned project
Python 3.15 reached its second release candidate, with the final release scheduled for 1 October 2026. The release notes list lazy imports, the Tachyon profiler and a built-in frozendict. UTF-8 also becomes the default encoding for open(), which could change how a CSV file reads on Windows. NumPy has prebuilt wheels for 3.15. pandas and scikit-learn wheels were still pending, so the advice is to stay on 3.14 for this course. Test 3.15 in a throwaway environment, and upgrade the lockfile only once those wheels exist. After upgrading, confirm the baseline still prints 11.35 percent.
Open the notebook from last time. You have two things loaded. X holds the pixels and y holds the labels. Nothing has learned anything yet. This chapter answers one question: how do you know whether a model is any good? The answer has two halves. A score means nothing until you compare it with a floor, and that floor has to be measured on data the model never saw. So before any clever model, we build the dumbest one possible and score it honestly.
Fitting a model that knows nothing
Start with NumPy, because NumPy is the shape of your data. Type X dot shape and run the cell. It prints seventy thousand, seven hundred eighty-four, inside parentheses. That pair is a description of a grid. There are seventy thousand rows and seven hundred eighty-four columns. Each row is one image. Each column is one pixel position, the same position in every image. Column zero is always the top-left corner, and column seven hundred eighty-three is always the bottom-right. This is the usual layout for tabular machine learning: one row per example, one column per feature. Almost every library you meet in this course expects it.
Now type y dot shape. It prints seventy thousand followed by a comma, all alone in parentheses. That trailing comma means a flat list with one entry per row of X. Row five of X pairs with entry five of y. That pairing is the whole supervised setup in a single line of output.
Next, ask about dtype, which is short for data type. X dot dtype tells you how each number is stored. Depending on how the loader returned it, you may see a floating-point type or a small integer type. Either way, the values run from zero to two hundred fifty-five, one brightness reading per pixel. The dtype matters for two practical reasons. It decides how much memory the array takes up, and it decides which arithmetic is safe. Small integers can overflow when you add many of them together, and floats cannot. Now check y dot dtype. You will see an object type. That is NumPy's way of saying these are text strings, the character "5" rather than the number five. You already knew the labels were strings. This is how you confirm it from the array itself, without trusting your memory.
Slicing is how you take pieces of an array. X with square brackets and colon sixty thousand means every row from the start up to, but not including, row sixty thousand. Check its shape and you get sixty thousand by seven hundred eighty-four. X with square brackets and sixty thousand colon means row sixty thousand through to the end. Its shape is ten thousand by seven hundred eighty-four. Add the two together and you get seventy thousand. Nothing is lost and nothing overlaps, because the stopping point of one slice is the starting point of the other. Do the same thing to y. Then you have four names: X_train, X_test, y_train and y_test.
Reshape turns one row back into a picture. Take X_train with square brackets and zero to get the first row, a flat line of seven hundred eighty-four numbers. Then call reshape with twenty-eight and twenty-eight. Twenty-eight times twenty-eight is seven hundred eighty-four, so the numbers fit exactly into a square, filled one row at a time. Print it and you see a block made mostly of zeros, with a curved band of large numbers in the middle. That band is the stroke of a handwritten five, which matches its label. Reshape copies nothing. It only changes how the same numbers are read. If you ask for a shape whose total does not equal seven hundred eighty-four, say twenty-eight by twenty-seven, NumPy refuses with an error. That refusal is useful, because it catches mistakes about what your data is.
Now move to pandas, the library for labeled tables. Here you use it for looking at the data. Wrap the training labels in a Series, a one-column labeled list. You write pd dot Series with y_train inside the parentheses. Then call value_counts on it. You get ten rows, one per digit, sorted from most common to least. The ones come first, at a little under seven thousand. The fives come last, at a little over fifty-four hundred. Call value_counts again with normalize equals True and you get proportions instead of counts. The ones sit near eleven percent. The others spread between about nine and eleven. So the classes are close to balanced, but not perfectly. Remember that eleven percent, because it comes back very soon.
Next, build a small DataFrame, which is pandas' word for a table with named columns. Give it one row per pixel and three columns: the mean of each column of X_train, the maximum, and the fraction of images where that pixel is zero. NumPy computes each of these across the rows in one line. Mean with axis equals zero means "average down each column." Hand the three results to pd dot DataFrame. Sort the table by the zero fraction and look at both ends. At the top are the border pixels, especially the corners. Their fraction is one, or very near it, and their mean and maximum are zero. In the whole training set, almost no one wrote in the corners. At the bottom are pixels near the center, which are inked in a large share of images.
That table teaches you something real. Some of the seven hundred eighty-four features carry almost no information, because a feature that is always zero cannot tell a three from an eight. That fact will matter when we talk about features and dimensionality reduction. There is also a common mistake to avoid here. It is tempting to think this pandas table is what the model eats. It is not. The model will train on X_train, the NumPy array, untouched. The DataFrame is a lens for you, the person inspecting the data. You build it, read it, and set it aside.
Before you fit anything, make a promise. The test slice is set aside now, before any modeling choice, and it stays shut. Here is why that matters. Each time you look at test results and then change something, you have quietly used the test set to choose. Change the model, check the test score, change a setting, check again. After twenty rounds, the configuration you keep is partly the one that happened to suit those particular ten thousand images. The final number then measures how well you tuned to the test set, not how well the model handles new handwriting. That small amount of peeking is enough to cost the number its honesty. So inspection uses training data, and choices use training data. The test set is scored once, at the end, as our learning-problem statement already requires.
You may wonder why taking the last ten thousand rows is safe, since slicing by position usually is not. It works for MNIST because the dataset was built with this split in mind. The first sixty thousand and last ten thousand were assembled so that each part has a mix of writers and digits. Most datasets offer no such guarantee. A file might be sorted by date, by customer or by label. If it were sorted by label, the last ten thousand rows might be nothing but nines. The general tool is a random shuffle split: scikit-learn's train_test_split. You give it X and y, a test_size such as zero point two for twenty percent, and a random_state, which is a seed so the same shuffle comes out every time. Passing stratify equals y keeps the class proportions the same on both sides. Shuffling is turned on by default, and stratifying requires it.
A random split also has a trap of its own, called leakage. Suppose one writer contributed fifty digits. A random shuffle can put some of them in training and some in test. Then the model is partly being tested on handwriting it has already seen, and the score comes out too high. The same thing happens with duplicate rows, or with several records from one customer. The fix is to split by the unit that has to stay separate, such as the writer, the patient or the time period. That is a choice about the problem, not a setting you can leave at its default. For MNIST, we keep the standard split.
Now the first fit. scikit-learn provides DummyClassifier in its dummy module. It is a model that ignores the pixels completely. It looks only at the labels. Create it with strategy equals most_frequent. Call fit with X_train and y_train. During fit, it counts the labels and stores which classes exist and how often each appears. It reads X only to check its size, never its values. It learns that "1" is the most common label. Call predict on X_test and you get ten thousand predictions, every one of them the string "1". Call score with X_test and y_test and you get the mean accuracy.
Before trusting that score, work out accuracy by hand on a tiny case. Say there are five test images with true labels 1, 7, 1, 3 and 0. The dummy predicts 1, 1, 1, 1, 1. It is right on the first and third, so two correct out of five, which is forty percent. Accuracy is just the number correct divided by the total. Now scale up. The test set contains one thousand one hundred thirty-five ones. Every one of those is correct, and every other image is wrong. One thousand one hundred thirty-five divided by ten thousand is about eleven point three five percent. That is the roughly eleven percent we wrote into the learning-problem statement, and now it has actually been measured.
Then look behind the single figure. Ask, for each true digit, what fraction the model got right. You can compare y_test with the predictions in a short loop, or put both into a pandas DataFrame and group by the true label. For ones the answer is one hundred percent. For zero, two, three and all the rest, it is zero. The overall figure of eleven percent sounds like a weak model that tries everything a little. In fact the model is perfect on one class and blind to the other nine. A single accuracy number averages that pattern away. A table with true labels down one side and predicted labels across the top would show it at a glance. That table is called a confusion matrix, and it is coming soon.
Here is an exercise, and the prediction is the important part. DummyClassifier has two other strategies that use chance. Stratified guesses at random in proportion to how common each class was in training. Uniform guesses each of the ten digits with equal probability. Both take a random_state, so set one. Before you run either, write down the accuracy you expect. Then reason it through. With uniform, each guess has a one-in-ten chance of matching, whatever the true digit is. With stratified, you are right about as often as a random guess with the training frequencies happens to match a random test label with the test frequencies. When the classes are close to balanced, both land near ten percent. Then run them and compare. Your numbers should come out close to ten percent, but not exactly, because a random guesser drifts a little from run to run. Also check per-class correctness. Both spread their errors evenly at about ten percent per digit, while most_frequent puts all its successes on one digit. Notice that most_frequent wins on accuracy and is the most lopsided of the three. That gap is the point of the exercise.
The floor now stands at about eleven point three five percent, measured on images the model never saw. Linear and logistic regression come next. They will be the first models that actually read the pixels, and their test accuracy only means something next to this number.
Python 3.15 and your pinned project
One news item matters to the workspace you just used. Python 3.15 reached its second release candidate on the first of September. Its final release is scheduled for the first of October, 2026. A release candidate is a nearly finished version made available for testing. Your project, built on 3.14, does not change on its own. The requires-python line in your project file sets the range of Python versions the project accepts, and uv.lock holds exact package versions. Nothing moves until you move it.
The release includes lazy imports to speed up startup, a new low-overhead profiler called Tachyon, and a built-in frozendict. One change could break old code: UTF-8 becomes the default text encoding for open() on every platform. A notebook that reads a CSV file without naming an encoding could behave differently on Windows.
The more practical concern is compiled packages. uv can already install 3.15 as a preview with uv python install three point fifteen, and NumPy has wheels ready, meaning prebuilt binary packages. Wheels for pandas and scikit-learn were still pending at the time of the release candidate, and they usually arrive after the final release. Until they arrive, a 3.15 environment may fail to install those packages or try to build them from source.
So stay on 3.14 for this course for now. It is still receiving bug fixes. If you want to test early, run your notebook's code under 3.15 in a throwaway environment and leave the lockfile alone. Once pandas and scikit-learn publish wheels, raise requires-python, run uv lock with the upgrade flag, and rerun from a clean sync. Check that the dummy baseline still prints the same eleven point three five percent.
