Want to understand how artificial intelligence works beyond calling a library function? One way to learn is to build a small neural network yourself.
In this tutorial, we will build a neural network from scratch using pure Python. We will not use TensorFlow, PyTorch, NumPy, or scikit-learn. Everything, including the forward pass, loss calculation, backpropagation, and weight updates, will be implemented using Python's standard library.
Our project will predict whether a student passes an exam based on the number of hours they studied. We will train the model on a small dataset and then test it on examples it has not seen during training.
I also ran the program in VS Code and included the actual output so you can compare your results with mine.
Table of Contents
- What We Are Building
- Set Up the Python Project
- Prepare the Dataset
- Build a Neural Network From Scratch
- Train the Neural Network
- Test the Model and Check the Results
- How the Neural Network Works
- Limitations and Experiments
- Final Thoughts
1. What We Are Building
Imagine we have a dataset containing students' study hours and exam results.
For example:
| Study hours | Exam result |
|---|---|
| 1 | Fail |
| 2 | Fail |
| 3 | Fail |
| 5 | Pass |
| 6 | Pass |
| 8 | Pass |
We could write a simple rule that predicts a pass whenever someone studies for at least five hours. However, that would be a rule we decided on ourselves.
Instead, we will create a neural network that learns from examples by adjusting numerical values called weights and biases.
The training process follows this general sequence:
Training Data
|
v
Make a Prediction
|
v
Calculate the Error
|
v
Update Weights and Biases
|
v
Repeat Training
|
v
Test on Unseen Examples
The network does not rewrite its Python code while learning. It changes the numerical parameters used to calculate its predictions.
2. Set Up the Python Project
You need Python 3 and a code editor. I used VS Code to run this project.
Create a folder containing these two files:
python-ai-from-scratch/
├── dataset.csv
└── ai.py
Check your Python installation in the terminal:
python --version
We will use three built-in Python modules:
-
csvto read the dataset. -
mathto perform mathematical calculations. -
randomto initialize weights and shuffle training examples.
You do not need to install any third-party packages.
3. Prepare the Dataset
Create a file named dataset.csv and add the following data:
study_hours,passed,split
1,0,train
2,0,train
2.5,0,train
3,0,train
3.5,0,test
4,0,train
4.5,1,train
5,1,train
5.5,1,test
6,1,train
7,1,train
8,1,test
The dataset contains three columns.
-
study_hoursrepresents the number of hours a student studied. -
passedcontains the expected result. A value of0means fail, while1means pass. -
splitidentifies whether a row belongs to the training set or the test set.
The model will learn from the training rows. The three rows marked test will be kept separate until evaluation.
These are synthetic examples created for this tutorial. They are not real student records, and they should not be used to draw conclusions about actual exam performance.
4. Build a Neural Network From Scratch
Now open ai.py and add the complete implementation below.
import csv
import math
import random
random.seed(42)
def sigmoid(value):
value = max(-500, min(500, value))
return 1 / (1 + math.exp(-value))
def leaky_relu(value):
# Keep a small gradient when the input is negative.
return value if value > 0 else 0.01 * value
def load_dataset(filename):
train_features = []
train_labels = []
test_features = []
test_labels = []
with open(filename, newline="") as file:
reader = csv.DictReader(file)
for row in reader:
hours = float(row["study_hours"])
# Scale study hours to a range from 0 to 1.
feature = [hours / 8]
label = int(row["passed"])
if row["split"] == "train":
train_features.append(feature)
train_labels.append(label)
else:
test_features.append(feature)
test_labels.append(label)
return (
train_features,
train_labels,
test_features,
test_labels,
)
class NeuralNetwork:
def __init__(self):
# Four neurons in the hidden layer.
self.hidden_weights = [
[random.uniform(-1, 1)]
for _ in range(4)
]
self.hidden_biases = [0.0] * 4
# One output neuron.
self.output_weights = [
random.uniform(-1, 1)
for _ in range(4)
]
self.output_bias = 0.0
def forward(self, inputs):
hidden = []
for weights, bias in zip(
self.hidden_weights,
self.hidden_biases,
):
value = sum(
weight * item
for weight, item in zip(weights, inputs)
) + bias
hidden.append(leaky_relu(value))
output_value = sum(
weight * value
for weight, value in zip(
self.output_weights, hidden
)
) + self.output_bias
return hidden, sigmoid(output_value)
def predict(self, inputs):
return self.forward(inputs)[1]
def calculate_loss(self, features, labels):
total_loss = 0.0
for inputs, target in zip(features, labels):
prediction = self.predict(inputs)
# Avoid taking the logarithm of zero.
p = min(
max(prediction, 1e-7),
1 - 1e-7,
)
total_loss -= (
target * math.log(p)
+ (1 - target) * math.log(1 - p)
)
return total_loss / len(features)
def train(self, features, labels, epochs=3000, rate=0.1):
for epoch in range(epochs):
# Shuffle training examples for each epoch.
order = list(range(len(features)))
random.shuffle(order)
for index in order:
inputs = features[index]
target = labels[index]
hidden, prediction = self.forward(inputs)
# Calculate the error at the output layer.
output_error = prediction - target
# Preserve the old weights for backpropagation.
old_output_weights = self.output_weights[:]
for i in range(len(self.output_weights)):
self.output_weights[i] -= (
rate * output_error * hidden[i]
)
self.output_bias -= rate * output_error
# Propagate the error through the hidden layer.
for i in range(len(self.hidden_weights)):
activation_gradient = (
1 if hidden[i] > 0 else 0.01
)
hidden_error = (
output_error
* old_output_weights[i]
* activation_gradient
)
self.hidden_weights[i][0] -= (
rate * hidden_error * inputs[0]
)
self.hidden_biases[i] -= (
rate * hidden_error
)
# Calculate loss after completing the epoch.
if epoch % 500 == 0:
loss = self.calculate_loss(features, labels)
print(
f"Epoch {epoch}: "
f"Training loss = {loss:.4f}"
)
if __name__ == "__main__":
(
train_features,
train_labels,
test_features,
test_labels,
) = load_dataset("dataset.csv")
model = NeuralNetwork()
model.train(train_features, train_labels)
print("\nTest results:")
correct = 0
for inputs, actual in zip(test_features, test_labels):
probability = model.predict(inputs)
predicted = int(probability >= 0.5)
if predicted == actual:
correct += 1
print(
f"Study hours: {inputs[0] * 8:g}, "
f"Actual: {actual}, "
f"Predicted: {predicted}, "
f"Score: {probability:.4f}"
)
accuracy = correct / len(test_labels)
print(
f"\nTest accuracy: {accuracy:.0%} "
f"({correct}/{len(test_labels)})"
)
Save the file and run the project from your terminal:
python ai.py
Here is what the main methods do:
-
load_dataset()reads the CSV file and separates training data from test data. -
forward()calculates the network's output. -
predict()returns the output score for an input. -
calculate_loss()measures the model's prediction error. -
train()updates the network's weights and biases.
Each method has a specific responsibility, which makes the code easier to understand and modify.
5. Train the Neural Network
Training means adjusting the network's parameters so that its predictions become closer to the expected answers.
The model starts with randomly initialized weights. During training, it processes examples, calculates prediction errors, and updates its parameters.
Consider this line:
output_error = prediction - target
If the expected result is 1 but the model predicts 0.3, the error is negative. The training update uses this information to move the parameters in a direction intended to reduce the loss.
Two important concepts are involved here.
Backpropagation calculates how the prediction error relates to parameters in the network.
Gradient descent uses those calculations to update the parameters.
The learning rate controls the size of each update:
model.train(
train_features,
train_labels,
epochs=3000,
rate=0.1,
)
A learning rate that is too large can make training unstable. A rate that is too small can make training progress slowly.
The code trains for 3,000 epochs and prints the training loss every 500 epochs. The loss is calculated after the updates for that epoch have finished.
6. Test the Model and Check the Results
After training, the program evaluates the three examples that were excluded from the training set.
For each example, it prints the study hours, actual label, predicted label, and output score.
I ran this implementation in VS Code. Here is the actual output from my run:
Epoch 0: Training loss = 0.6799
Epoch 500: Training loss = 0.0387
Epoch 1000: Training loss = 0.0099
Epoch 1500: Training loss = 0.0053
Epoch 2000: Training loss = 0.0035
Epoch 2500: Training loss = 0.0026
Test results:
Study hours: 3.5, Actual: 0, Predicted: 0, Score: 0.0009
Study hours: 5.5, Actual: 1, Predicted: 1, Score: 1.0000
Study hours: 8, Actual: 1, Predicted: 1, Score: 1.0000
Test accuracy: 100% (3/3)
The training loss decreased from 0.6799 at epoch 0 to 0.0026 at epoch 2500. This shows that the model became much better at fitting the training examples.
The three test predictions were also correct:
| Study hours | Actual result | Predicted result |
|---|---|---|
| 3.5 | Fail | Fail |
| 5.5 | Pass | Pass |
| 8 | Pass | Pass |
The program reports 100% test accuracy because it correctly classified all three test examples.
Accuracy is calculated using this formula:
Accuracy (%) = (Correct Predictions / Total Test Examples) * 100
For this run:
Accuracy (%) = (3 / 3) * 100
Accuracy (%) = 100%
The score values are displayed to four decimal places. Therefore, a printed score of 1.0000 may be a value very close to 1 rather than exactly 1.
There is also an important limitation to keep in mind. The test set contains only three examples. If the model got one prediction wrong, its accuracy would fall to 66.67%.
These results show that the program ran and classified this small test set correctly. They do not prove that it will perform equally well on new or real-world data.
7. How the Neural Network Works
The implementation uses a few mathematical operations to turn input data into predictions.
Forward propagation
Forward propagation passes the input through the hidden layer and then the output layer. Each neuron combines its inputs with weights and adds a bias.
The final output is a score between zero and one.
Leaky ReLU
The hidden neurons use Leaky ReLU as their activation function.
For a positive input, the function returns the input unchanged. For a negative input, it returns a small negative value instead of zero.
This small slope allows gradients to continue flowing through negative activations, reducing the risk of permanently inactive neurons.
Sigmoid
The output neuron uses the sigmoid function:
def sigmoid(value):
value = max(-500, min(500, value))
return 1 / (1 + math.exp(-value))
Sigmoid converts a number into a value between zero and one. The model uses 0.5 as its classification threshold:
- A score below
0.5becomes class0. - A score equal to or above
0.5becomes class1.
Loss calculation
The network uses binary cross-entropy to measure how far its predictions are from the expected labels.
A lower loss generally means the model fits the training examples better. However, a low training loss does not guarantee good performance on unseen data.
Backpropagation and gradient descent
Backpropagation calculates the gradients needed to update the network's parameters. Gradient descent then uses those gradients and the learning rate to change the weights and biases.
Together, these steps allow the model to learn from examples without manually writing a rule for every possible input.
8. Limitations and Experiments
This project is intentionally small. It uses one input feature, four hidden neurons, nine training examples, and three test examples.
The dataset is synthetic, and the test set is too small to provide strong evidence of generalization. The results should be treated as a demonstration of how a neural network works, not as a reliable exam prediction system.
Here are a few experiments you can try:
- Change the learning rate from
0.1to0.01and compare the training loss. - Increase the number of hidden neurons and see how the training results change.
- Add more examples to the dataset and reserve a larger test set.
- Add more input features, if you have suitable data.
- Save the trained weights and biases so you can reuse the model without training it again.
Change one setting at a time and record what happens. That makes it easier to understand which change caused a difference.
This project is also very different from a large language model. Modern language models use much larger networks, extensive training datasets, and substantial computing resources. Still, the basic ideas of parameters, predictions, loss, and optimization are useful starting points for understanding them.
9. Final Thoughts
We have built a small neural network from scratch using Python's standard library. It reads a dataset, makes predictions, calculates errors, updates its parameters, and evaluates its predictions on separate test examples.
The most useful part of this exercise is being able to inspect the calculations instead of treating model training as a black box.
Start by changing the learning rate, run the program again, and compare the results. Once you understand this version, you can move on to larger datasets and more advanced neural network architectures.
Top comments (1)
Official Platform Update
Security protocols have been updated for all developer accounts.