DEV Community

Cover image for How to Build a Neural Network From Scratch in Python
Arafat Hossain Ar
Arafat Hossain Ar Subscriber

Posted on

How to Build a Neural Network From Scratch in Python

Want to understand how artificial intelligence works beyond calling a library function? One way to learn is to build a small neural network yourself.

In this tutorial, we will build a neural network from scratch using pure Python. We will not use TensorFlow, PyTorch, NumPy, or scikit-learn. Everything, including the forward pass, loss calculation, backpropagation, and weight updates, will be implemented using Python's standard library.

Our project will predict whether a student passes an exam based on the number of hours they studied. We will train the model on a small dataset and then test it on examples it has not seen during training.

I also ran the program in VS Code and included the actual output so you can compare your results with mine.

Table of Contents

  1. What We Are Building
  2. Set Up the Python Project
  3. Prepare the Dataset
  4. Build a Neural Network From Scratch
  5. Train the Neural Network
  6. Test the Model and Check the Results
  7. How the Neural Network Works
  8. Limitations and Experiments
  9. Final Thoughts

1. What We Are Building

Imagine we have a dataset containing students' study hours and exam results.

For example:

Study hours Exam result
1 Fail
2 Fail
3 Fail
5 Pass
6 Pass
8 Pass

We could write a simple rule that predicts a pass whenever someone studies for at least five hours. However, that would be a rule we decided on ourselves.

Instead, we will create a neural network that learns from examples by adjusting numerical values called weights and biases.

The training process follows this general sequence:

Training Data
     |
     v
Make a Prediction
     |
     v
Calculate the Error
     |
     v
Update Weights and Biases
     |
     v
Repeat Training
     |
     v
Test on Unseen Examples
Enter fullscreen mode Exit fullscreen mode

The network does not rewrite its Python code while learning. It changes the numerical parameters used to calculate its predictions.

2. Set Up the Python Project

You need Python 3 and a code editor. I used VS Code to run this project.

Create a folder containing these two files:

python-ai-from-scratch/
├── dataset.csv
└── ai.py
Enter fullscreen mode Exit fullscreen mode

Check your Python installation in the terminal:

python --version
Enter fullscreen mode Exit fullscreen mode

We will use three built-in Python modules:

  • csv to read the dataset.
  • math to perform mathematical calculations.
  • random to initialize weights and shuffle training examples.

You do not need to install any third-party packages.

3. Prepare the Dataset

Create a file named dataset.csv and add the following data:

study_hours,passed,split
1,0,train
2,0,train
2.5,0,train
3,0,train
3.5,0,test
4,0,train
4.5,1,train
5,1,train
5.5,1,test
6,1,train
7,1,train
8,1,test
Enter fullscreen mode Exit fullscreen mode

The dataset contains three columns.

  • study_hours represents the number of hours a student studied.
  • passed contains the expected result. A value of 0 means fail, while 1 means pass.
  • split identifies whether a row belongs to the training set or the test set.

The model will learn from the training rows. The three rows marked test will be kept separate until evaluation.

These are synthetic examples created for this tutorial. They are not real student records, and they should not be used to draw conclusions about actual exam performance.

4. Build a Neural Network From Scratch

Now open ai.py and add the complete implementation below.

import csv
import math
import random

random.seed(42)


def sigmoid(value):
    value = max(-500, min(500, value))
    return 1 / (1 + math.exp(-value))


def leaky_relu(value):
    # Keep a small gradient when the input is negative.
    return value if value > 0 else 0.01 * value


def load_dataset(filename):
    train_features = []
    train_labels = []
    test_features = []
    test_labels = []

    with open(filename, newline="") as file:
        reader = csv.DictReader(file)

        for row in reader:
            hours = float(row["study_hours"])

            # Scale study hours to a range from 0 to 1.
            feature = [hours / 8]
            label = int(row["passed"])

            if row["split"] == "train":
                train_features.append(feature)
                train_labels.append(label)
            else:
                test_features.append(feature)
                test_labels.append(label)

    return (
        train_features,
        train_labels,
        test_features,
        test_labels,
    )


class NeuralNetwork:
    def __init__(self):
        # Four neurons in the hidden layer.
        self.hidden_weights = [
            [random.uniform(-1, 1)]
            for _ in range(4)
        ]
        self.hidden_biases = [0.0] * 4

        # One output neuron.
        self.output_weights = [
            random.uniform(-1, 1)
            for _ in range(4)
        ]
        self.output_bias = 0.0

    def forward(self, inputs):
        hidden = []

        for weights, bias in zip(
            self.hidden_weights,
            self.hidden_biases,
        ):
            value = sum(
                weight * item
                for weight, item in zip(weights, inputs)
            ) + bias

            hidden.append(leaky_relu(value))

        output_value = sum(
            weight * value
            for weight, value in zip(
                self.output_weights, hidden
            )
        ) + self.output_bias

        return hidden, sigmoid(output_value)

    def predict(self, inputs):
        return self.forward(inputs)[1]

    def calculate_loss(self, features, labels):
        total_loss = 0.0

        for inputs, target in zip(features, labels):
            prediction = self.predict(inputs)

            # Avoid taking the logarithm of zero.
            p = min(
                max(prediction, 1e-7),
                1 - 1e-7,
            )

            total_loss -= (
                target * math.log(p)
                + (1 - target) * math.log(1 - p)
            )

        return total_loss / len(features)

    def train(self, features, labels, epochs=3000, rate=0.1):
        for epoch in range(epochs):
            # Shuffle training examples for each epoch.
            order = list(range(len(features)))
            random.shuffle(order)

            for index in order:
                inputs = features[index]
                target = labels[index]

                hidden, prediction = self.forward(inputs)

                # Calculate the error at the output layer.
                output_error = prediction - target

                # Preserve the old weights for backpropagation.
                old_output_weights = self.output_weights[:]

                for i in range(len(self.output_weights)):
                    self.output_weights[i] -= (
                        rate * output_error * hidden[i]
                    )

                self.output_bias -= rate * output_error

                # Propagate the error through the hidden layer.
                for i in range(len(self.hidden_weights)):
                    activation_gradient = (
                        1 if hidden[i] > 0 else 0.01
                    )

                    hidden_error = (
                        output_error
                        * old_output_weights[i]
                        * activation_gradient
                    )

                    self.hidden_weights[i][0] -= (
                        rate * hidden_error * inputs[0]
                    )

                    self.hidden_biases[i] -= (
                        rate * hidden_error
                    )

            # Calculate loss after completing the epoch.
            if epoch % 500 == 0:
                loss = self.calculate_loss(features, labels)

                print(
                    f"Epoch {epoch}: "
                    f"Training loss = {loss:.4f}"
                )


if __name__ == "__main__":
    (
        train_features,
        train_labels,
        test_features,
        test_labels,
    ) = load_dataset("dataset.csv")

    model = NeuralNetwork()

    model.train(train_features, train_labels)

    print("\nTest results:")

    correct = 0

    for inputs, actual in zip(test_features, test_labels):
        probability = model.predict(inputs)
        predicted = int(probability >= 0.5)

        if predicted == actual:
            correct += 1

        print(
            f"Study hours: {inputs[0] * 8:g}, "
            f"Actual: {actual}, "
            f"Predicted: {predicted}, "
            f"Score: {probability:.4f}"
        )

    accuracy = correct / len(test_labels)

    print(
        f"\nTest accuracy: {accuracy:.0%} "
        f"({correct}/{len(test_labels)})"
    )
Enter fullscreen mode Exit fullscreen mode

Save the file and run the project from your terminal:

python ai.py
Enter fullscreen mode Exit fullscreen mode

Here is what the main methods do:

  • load_dataset() reads the CSV file and separates training data from test data.
  • forward() calculates the network's output.
  • predict() returns the output score for an input.
  • calculate_loss() measures the model's prediction error.
  • train() updates the network's weights and biases.

Each method has a specific responsibility, which makes the code easier to understand and modify.

5. Train the Neural Network

Training means adjusting the network's parameters so that its predictions become closer to the expected answers.

The model starts with randomly initialized weights. During training, it processes examples, calculates prediction errors, and updates its parameters.

Consider this line:

output_error = prediction - target
Enter fullscreen mode Exit fullscreen mode

If the expected result is 1 but the model predicts 0.3, the error is negative. The training update uses this information to move the parameters in a direction intended to reduce the loss.

Two important concepts are involved here.

Backpropagation calculates how the prediction error relates to parameters in the network.

Gradient descent uses those calculations to update the parameters.

The learning rate controls the size of each update:

model.train(
    train_features,
    train_labels,
    epochs=3000,
    rate=0.1,
)
Enter fullscreen mode Exit fullscreen mode

A learning rate that is too large can make training unstable. A rate that is too small can make training progress slowly.

The code trains for 3,000 epochs and prints the training loss every 500 epochs. The loss is calculated after the updates for that epoch have finished.

6. Test the Model and Check the Results

After training, the program evaluates the three examples that were excluded from the training set.

For each example, it prints the study hours, actual label, predicted label, and output score.

I ran this implementation in VS Code. Here is the actual output from my run:

Epoch 0: Training loss = 0.6799
Epoch 500: Training loss = 0.0387
Epoch 1000: Training loss = 0.0099
Epoch 1500: Training loss = 0.0053
Epoch 2000: Training loss = 0.0035
Epoch 2500: Training loss = 0.0026

Test results:
Study hours: 3.5, Actual: 0, Predicted: 0, Score: 0.0009
Study hours: 5.5, Actual: 1, Predicted: 1, Score: 1.0000
Study hours: 8, Actual: 1, Predicted: 1, Score: 1.0000

Test accuracy: 100% (3/3)
Enter fullscreen mode Exit fullscreen mode

The training loss decreased from 0.6799 at epoch 0 to 0.0026 at epoch 2500. This shows that the model became much better at fitting the training examples.

The three test predictions were also correct:

Study hours Actual result Predicted result
3.5 Fail Fail
5.5 Pass Pass
8 Pass Pass

The program reports 100% test accuracy because it correctly classified all three test examples.

Accuracy is calculated using this formula:

Accuracy (%) = (Correct Predictions / Total Test Examples) * 100
Enter fullscreen mode Exit fullscreen mode

For this run:

Accuracy (%) = (3 / 3) * 100
Accuracy (%) = 100%
Enter fullscreen mode Exit fullscreen mode

The score values are displayed to four decimal places. Therefore, a printed score of 1.0000 may be a value very close to 1 rather than exactly 1.

There is also an important limitation to keep in mind. The test set contains only three examples. If the model got one prediction wrong, its accuracy would fall to 66.67%.

These results show that the program ran and classified this small test set correctly. They do not prove that it will perform equally well on new or real-world data.

7. How the Neural Network Works

The implementation uses a few mathematical operations to turn input data into predictions.

Forward propagation

Forward propagation passes the input through the hidden layer and then the output layer. Each neuron combines its inputs with weights and adds a bias.

The final output is a score between zero and one.

Leaky ReLU

The hidden neurons use Leaky ReLU as their activation function.

For a positive input, the function returns the input unchanged. For a negative input, it returns a small negative value instead of zero.

This small slope allows gradients to continue flowing through negative activations, reducing the risk of permanently inactive neurons.

Sigmoid

The output neuron uses the sigmoid function:

def sigmoid(value):
    value = max(-500, min(500, value))
    return 1 / (1 + math.exp(-value))
Enter fullscreen mode Exit fullscreen mode

Sigmoid converts a number into a value between zero and one. The model uses 0.5 as its classification threshold:

  • A score below 0.5 becomes class 0.
  • A score equal to or above 0.5 becomes class 1.

Loss calculation

The network uses binary cross-entropy to measure how far its predictions are from the expected labels.

A lower loss generally means the model fits the training examples better. However, a low training loss does not guarantee good performance on unseen data.

Backpropagation and gradient descent

Backpropagation calculates the gradients needed to update the network's parameters. Gradient descent then uses those gradients and the learning rate to change the weights and biases.

Together, these steps allow the model to learn from examples without manually writing a rule for every possible input.

8. Limitations and Experiments

This project is intentionally small. It uses one input feature, four hidden neurons, nine training examples, and three test examples.

The dataset is synthetic, and the test set is too small to provide strong evidence of generalization. The results should be treated as a demonstration of how a neural network works, not as a reliable exam prediction system.

Here are a few experiments you can try:

  1. Change the learning rate from 0.1 to 0.01 and compare the training loss.
  2. Increase the number of hidden neurons and see how the training results change.
  3. Add more examples to the dataset and reserve a larger test set.
  4. Add more input features, if you have suitable data.
  5. Save the trained weights and biases so you can reuse the model without training it again.

Change one setting at a time and record what happens. That makes it easier to understand which change caused a difference.

This project is also very different from a large language model. Modern language models use much larger networks, extensive training datasets, and substantial computing resources. Still, the basic ideas of parameters, predictions, loss, and optimization are useful starting points for understanding them.

9. Final Thoughts

We have built a small neural network from scratch using Python's standard library. It reads a dataset, makes predictions, calculates errors, updates its parameters, and evaluates its predictions on separate test examples.

The most useful part of this exercise is being able to inspect the calculations instead of treating model training as a black box.

Start by changing the learning rate, run the program again, and compare the results. Once you understand this version, you can move on to larger datasets and more advanced neural network architectures.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

Official Platform Update

Security protocols have been updated for all developer accounts.

  • tr.ee/dev-to