0.1 Building Vanilla NN From Scratch
For our toy examples, we will create two things. The first one is technically not a neural network as it does not contain conventional neurons and structures. The second one is an actual neural net that is very simple. In this section, I will share the entire code first, then provide explanations.
0.1.1 Simple Stochastic Gradient Descent-Based
Learning Engine
What we will make now is a simple stochastic gradient descent-based learning engine. This isn’t technically a neural network, as you will see why in the code, but I decided to include it here as making similar examples really helped me understand backprop with code.
Our task with this learning engine is simple. We are going to predict \(e^x\) for limited range with concatenated Maclaurin Series. If you recall from your calculus courses, the following equation holds by Taylor series. \[ e^x = \sum _{n = 0}^\infty \frac {1}{n!} x^n = 1 + x + \frac {x^2}{2!} + \frac {x^3}{3!} + \cdots \] Here, for limited \(x\), we will be approximating \(e^x\) with the following polynomial where \(b\) is the bias and \(w_i\) are the weights. \[ e^x \approx b + w_1x + w_2x^2 + w_3x^3 + w_4x^4 + w_5x^5 + w_6x^6 + w_7x^7 + w_8x^8 \] We will get back to this on why our \(x\) should be limited, but let’s see our structure first.
We will be “training” the engine with \(200\) examples from \(-1.00\) to \(0.99\) inclusive. Then, we will be testing it with \(50\) validations from \(1.00\) to \(1.49\) inclusive. This backprop machine is stochastic as it will adhere to the following structure.
As you can see, we will be using MSE for our loss function, and we do not need activation functions like sigmoid and tanh because we don’t need one in our simple example. Okay now here is the breakdown of the entire notebook available here.
Detailed Explanation
Now let’s take a look at individual cells.
Here, I tried to keep it as minimal as possible. Instead of NumPy, I decided to
use the math module for all math stuffs like math.exp(). We need random
as it would make our lives easier than to manually set the eight weights and
the bias. Matplotlib of course since we want to see how it learns for limited
range of input.
1train_x = [i/100 for i in range(-100, 100)] 2train_y = [math.exp(i/100) for i in range(-100, 100)] 3test_x = [i/100 for i in range(100, 150)] 4test_y = [math.exp(i/100) for i in range(100, 150)]
Setting up the dataset for our engine is also quite simple. Instead of writing
every single number, I just decided to loop through them and store them in
an array, and that is train_x. The values for train_y is also quite
straightforward as it is \(e^x\) for corresponding \(x\). The same goes for the test set. I
tried to keep it \(80\%\) and \(20\%\) for training and validation, but that is not necessary and
it varies by models. However, I decided to use that here because that’s pretty
much what most people do when they first train. Now let’s get into the fun
stuffs.
1class SGDBLE: 2 def __init__(self): 3 random.seed(3141592) # So that we get the same result 4 self.ws = [random.uniform(-1, 1) for _ in range(8)] 5 self.b = random.random()
Now I decided to store the core functions in the class SGDBLE as we will be
using self to store attributes and objects. Here we set the specific seed so
that you will also get the same result when you run the code in somewhere.
Next we set weights and biases and stored them in self with our set
random values.
1def forward(self, x): 2 self.ypred = sum(w * x**k for w, k in zip(self.ws,range(1, 9))) + self.b 3 return self.ypred 4 5def loss(self, y, ypred): 6 self.loss_val = 0.5 * (y - ypred)**2 7 return self.loss_val
This is where we set some core functions. Forward is forward pass and it is somewhat straight forward as that is how we defined it. For the loss function, we are using MSE, i.e. \(\frac {1}{2} (y - \hat {y})^2\).
1def grads(self, x, y, ypred): 2 base_der = ypred - y # i.e. dL / dypred 3 self.grads_w = [base_der * x**i for i in range(1, 9)] 4 self.grad_b = base_der 5 6def update(self, lr): 7 for i in range(8): 8 self.ws[i] -= lr * self.grads_w[i] 9 self.b -= lr * self.grad_b
This is the core part of our engine. Let’s briefly look at the math and write the
loop out for some to see what’s actually going on. Consider the following
equation by the chain rule. \[ \frac {dL}{dw_i} = \frac {dL}{dy} \cdot \frac {dy}{dw_i} \] Notice that \(\frac {dL}{dy} = \frac {d}{dy} (\frac {1}{2} (y - \hat {y})^2) = y - \hat {y}\). This is base_der. Continuing with
the chain rule, we can write the weights and the bias as the following.
\begin{align*} \frac {dL}{dw_1} &= \frac {dL}{dy} \cdot \frac {dy}{dw_1} = \text {base\_der} \cdot x \\ \frac {dL}{dw_2} &= \frac {dL}{dy} \cdot \frac {dy}{dw_2} = \text {base\_der} \cdot x^2 \\ \frac {dL}{db} &= \frac {dL}{dy} \cdot \frac {dy}{db} = \text {base\_der} \cdot 1 \end{align*}
Now the update function. Recall \(w_2 = w_1 - \eta \frac {dL}{dw_1}\) and \(b_2 = b_1 - \eta \frac {dL}{db_1}\). That is exactly what we did, but for weights we have to loop through each element.
1def train(self, epochs, lr): 2 losses = [] 3 for epoch in range(epochs): 4 loss_total = 0 5 for i in range(len(train_x)): 6 ypred = self.forward(train_x[i]) 7 loss_total += self.loss(train_y[i], ypred) 8 self.grads(train_x[i], train_y[i], ypred) 9 self.update(lr) 10 loss_avg = loss_total / len(train_x) 11 losses.append(loss_avg) 12 13 if (epoch+1) % 100 == 0: 14 print(f"Epoch: {epoch+1}/{epochs}, Loss: {loss_avg}") 15 16 return losses
This is the training loop. Of course we don’t want to stop at a single epoch
especially in stochastic setting. You can ignore the losses in the beginning
as we only use them to graph. For each epoch, we want to set the total loss to
\(0\) as they will accumulate. Then, for all \(x\) values in our training set, we will
predict, calculate the loss, backprop, an update as shown. Note that we have
to set epoch and lr, which is the learning rate. Now we will take the
average of the total loss and print it every \(100\) epoch to see how well our engine
is learning.
1def test(self): 2 loss_total = 0 3 for i in range(0, len(test_x), 2): 4 ypred1 = self.forward(test_x[i]) 5 loss_total += self.loss(test_y[i], ypred1) 6 7 if i + 1 len(test_x): 8 ypred2 = self.forward(test_x[i + 1]) 9 loss_total += self.loss(test_y[i + 1], ypred2) 10 11 print(f"Test {i+1} Test {i+2}") 12 print(f"y: {test_y[i]} y: {test_y[i+1]}") 13 print(f"ypred: {ypred1} ypred: {ypred2}") 14 print() 15 else: 16 print(f"Test {i+1}") 17 print(f"y: {test_y[i]}") 18 print(f"ypred: {ypred1}") 19 20 loss_total /= len(test_x) 21 22def print_params(self): 23 coeff = [1 / math.factorial(i+1) for i in range(8)] 24 for i in range(8): 25 print(f"w{i+1}: {self.ws[i]}, Expected: {coeff[i]}") 26 print(f"b: {self.b}, Expected: 1") 27 28model = SGDBLE()
The test loop is also straight forward. It predicts, calculates the loss, and print them as shown. Also, because we know the coefficients for the Maclaurin Series, I decided to see if the parameters match with the actual coefficients. Finally, the last line is to be able use the model.
1def graph_loss(losses): 2 plt.plot(losses) 3 plt.xlabel("Epoch") 4 plt.ylabel("Loss") 5 plt.show() 6 7def graph_pred(): 8 xs = [i / 10 for i in range(-100, 200)] 9 true = [math.exp(x) for x in xs] 10 pred = [model.forward(x) for x in xs] 11 12 plt.plot(xs, true, label="e^x") 13 plt.plot(xs, pred, label="Prediction") 14 plt.yscale("log") 15 plt.legend() 16 plt.show()
This part is plotting with Matplotlib. Please refer to the second note if you
are confused! One thing though yscale is necessary as exponential function
will get too large to show.
This part is the actual training part. We set the epoch to \(1000\) and learning rate to \(0.005\). This learning rate might be too small for the start, but I realized that this gives slightly better results than \(0.001\) in this setting. Then we graph and show loss, parameters, and predictions.
This part is for testing. Now that we have seen the entire code and explanation for it, let’s now analyze our results.
Analyzing the Results
First, let’s take a look at the first result.
1losses = model.train(1000, 0.005) 2graph_loss(losses) 3model.print_params() 4graph_pred()
Epoch: 100/1000, Loss: 0.00013234985876534958 Epoch: 200/1000, Loss: 3.427320321762344e-05 Epoch: 300/1000, Loss: 2.4492707827970274e-05 Epoch: 400/1000, Loss: 2.2669862503550263e-05 Epoch: 500/1000, Loss: 2.172205317266906e-05 Epoch: 600/1000, Loss: 2.0984248749352728e-05 Epoch: 700/1000, Loss: 2.036538265870468e-05 Epoch: 800/1000, Loss: 1.9834658817269138e-05 Epoch: 900/1000, Loss: 1.937270927247297e-05 Epoch: 1000/1000, Loss: 1.8965386261105253e-05
w1: 0.9444454598562942, Expected: 1.0 w2: 0.49986306356857685, Expected: 0.5 w3: 0.620431907367758, Expected: 0.16666666666666666 w4: 0.21618569290559692, Expected: 0.041666666666666664 w5: -0.9321256839917001, Expected: 0.008333333333333333 w6: -0.48709437200615985, Expected: 0.001388888888888889 w7: 0.5594822789317027, Expected: 0.0001984126984126984 w8: 0.33002989302732855, Expected: 2.48015873015873e-05 b: 0.9983077779493585, Expected: 1
![]()
This is what we got after the first training. The loss looks pretty good actually. As you can see for each \(100^\text {th}\) epoch, the loss is decreasing well. From the graph, you can see the sudden drop in losses; however, we can’t necessarily say that this is a grokking phenomena as it might just be model understanding the pattern from underfitting.
Looking at the weights, some are good and some are bad. We will get back to this later, but this is also one of the reason why we cannot scale this engine. Similarly, we can see that our polynomial fits very well for the small range, but not well as the input increases.
1model.test()
Test 1 Test 2 y: 2.718281828459045 y: 2.7456010150169163 ypred: 2.7495260176087557 ypred: 2.7867851769675642 Test 3 Test 4 y: 2.7731947639642978 y: 2.801065834699079 ypred: 2.8257757734623574 ypred: 2.8666322590606814 Test 5 Test 6 y: 2.82921701435156 y: 2.857651118063164 ypred: 2.909497162739583 ypred: 2.954521430346852 Test 7 Test 8 y: 2.8863709892679585 y: 2.915379499976997 ypred: 3.001864774074873 ypred: 3.0516960317129964 Test 9 Test 10 y: 2.944679551065524 y: 2.9742740725630656 ypred: 3.1041935358457007 ypred: 3.159545493165112 Test 11 Test 12 y: 3.0041660239464334 y: 3.034358394435676 ypred: 3.2179503740678106 ypred: 3.2796173127071624 Test 13 Test 14 y: 3.0648542032930024 y: 3.095656500124711 ypred: 3.3447665176737535 ypred: 3.4136296934778425 Test 15 Test 16 y: 3.1267683651861553 y: 3.158192909689767 ypred: 3.4864504730090555 ypred: 3.5634848611499024 ... Test 45 Test 46 y: 4.220695816996552 y: 4.263114515168817 ypred: 9.347554553339133 ypred: 9.752958544358306 Test 47 Test 48 y: 4.305959528345206 y: 4.349235141062741 ypred: 10.179646340417854 ypred: 10.628580569357089 Test 49 Test 50 y: 4.392945680918757 y: 4.437095519003664 ypred: 11.100758361572902 ypred: 11.597212287422602
For the training set, I only listed some of them as you can refer to the previous pages. Notice that for first few tests, the predictions are not bad. However, for Test 5, the engine does not seem be effective and seems lost from Test 10. We could actually run our training loop multiple times to see what happens, and I will leave this to the readers. However, the engine does not seem to learn well.
Here is the question. Is the engine overfitting right now? If you said no, well... that is correct! Technically, we can see it as overfitting, but what I think is more accurate is to say that this is the limitation of the scalability of this engine and our method. When predicting, we only used nine terms. From the parameters, some were very close to the expectation and others were very bad. This is because we are using a polynomial with finite term to predict, whereas the real Maclaurin series uses infinite sum. Therefore, it is techinically an overfitting to certain point in a way that our engine tries to fit our polynomial in limited range as shown in the graph, but this too could improve if we have the computational ability to use infinite sum. Therefore, it may be more accurate to say that it is the limitation of the engine and scalability that causes high loss in testing.
We just completed our first model! Although it is very limited in terms of scalability and usability, I hope this example was very helpful in using backpropagation in code. For our next example, we will build an actual simple neural net to predict a nonlinear function.
0.1.2 Vanilla Neural Network
Now that we have created a simple backprop engine, let’s increase the number of hidden layers to predict \(x_1 x_2\) for inputs \(x_1\) and \(x_2\). I thought it would be great to introduce nonlinearity to our problem, and it works! Let’s started with detailed explanations of the entire notebook available here.
Detailed Explanation
Now let’s take a look at each block of the code to see their role in our vanilla neural net.
Here, I imported essential modules and libraries to do math, get a randomized value, and plot the loss.
1# Predicting x1 * x2 2def tanh(x): 3 return (math.exp(x) - math.exp(-x)) / (math.exp(x) + math.exp(-x)) 4 5def tanh_der(x): 6 return 1 - tanh(x)**2 7 8def loss(y, ypred): 9 return 0.5 * (y - ypred)**2
For this problem, I decided to use tanh for our activation function as sigmoid squishes the values to 0 and 1, which is not ideal for regression problem like ours and ReLU was not effective when I tried. Since \(\tanh (x) = \frac {e^x - e^{-x}}{e^x + e^{-x}}\) and \(\tanh '(x) = 1 - \tanh ^2(x)\), we did the same. For our loss function, I decided to use MSE for simplicity again.
1# Data 2data_x = [[i/10, j/10] for i in range(-10, 11) for j in range(-10, 11)] 3data_y = [(x1 * x2) for x1, x2 in data_x] 4 5data = list(zip(data_x, data_y)) 6random.seed(3141592) 7random.shuffle(data) 8 9split = int(0.8 * len(data)) 10train_data = data[:split] 11test_data = data[split:] 12 13train_x = [x for x, y in train_data] 14train_y = [y for x, y in train_data] 15test_x = [x for x, y in test_data] 16test_y = [y for x, y in test_data]
This part is for making our datasets. I decided to divide \(i\) and \(j\) by \(10\) because if I don’t, math overflow error will be resulted due to \(\tanh \) function.
This time, I decided to shuffle data in a way that we will get the same result when you try running this, and split the data into \(80\%\) for training and \(20\%\) for validating.
1class MLP: 2 def __init__(self, layers): 3 self.ws = [] 4 self.bs = [] 5 6 random.seed(3141592) 7 for i in range(len(layers) - 1): 8 layer_in = layers[i] 9 layer_out = layers[i + 1] 10 layer_w = [[random.uniform(-1, 1) for _ in range(layer_in)] 11 for _ in range(layer_out)] 12 layer_b = [random.uniform(-1, 1) for _ in range(layer_out)] 13 14 self.ws.append(layer_w) 15 self.bs.append(layer_b)
Again for our example, I decided to store our core functions in a class since
we want to reuse and store our attributes. When this class is first called, we
will create an empty array and for ws and bs and store them in self. Then,
for the randomized value, we will loop through each layers to store a
randomized value between \(-1\) and \(1\) for weights and biases. Notice that
unlike the bias which only loop through the neurons in the output
layer, we have to loop through the previous layer too for weights
since we have different weights for each previous neuron for each
output neuron. We will then append, or store, our layer to ws and
bs.
1def forward(self, x): 2 self.act = [x] 3 self.zs = [] 4 5 # Layers 6 for i in range(len(self.ws)): 7 prev = self.act[-1] 8 pre_layer = [] 9 act_layer = [] 10 11 # Neurons 12 for j in range(len(self.ws[i])): 13 pre = self.bs[i][j] 14 for k in range(len(self.ws[i][j])): 15 pre += self.ws[i][j][k] * prev[k] 16 17 pre_layer.append(pre) 18 19 if i != len(self.ws) - 1: 20 act_layer.append(tanh(pre)) 21 else: 22 act_layer.append(pre) 23 24 self.zs.append(pre_layer) 25 self.act.append(act_layer) 26 27 return self.act[-1][0]
This is the forward function. Here, I decided to create both self.act and
self.zs since we need both of them for backprop function. As the name
suggests, self.act is the value after \(\tanh \) is applied to self.zs. Moreover, I
stored the input in self.act in the beginning to start using them in the first
hidden layer.
Continuing, we create arrays for each layer for preactivation and activation.
Then, we will continue by looping through the neurons within each layer.
Our preactivation will start from the bias, and by following the definition,
we will add the products of weights and the previous activation. For the
if-else condition, I decided to do this because we don’t want \(\tanh \) for our final
output layer. Then, we will store the layers in self.act and self.zs for
latter use.
1def backprop(self, y): 2 self.grad_w = [[[0.0 for _ in range(len(self.ws[i][j]))] 3 for j in range(len(self.ws[i]))] 4 for i in range(len(self.ws))] 5 self.grad_b = [[0.0 for _ in range(len(self.bs[i]))] 6 for i in range(len(self.bs))] 7 ypred = self.act[-1][0] 8 delta = [0] * len(self.ws) 9 delta[-1] = [ypred - y] 10 11 # range(start, stop, step) 12 for k in range(len(self.ws) - 2, -1, -1): 13 delta[k] = [] 14 for i in range(len(self.ws[k])): 15 dlda = 0.0 16 for j in range(len(self.ws[k + 1])): 17 dlda += self.ws[k + 1][j][i] * delta[k + 1][j] 18 delta[k].append(dlda * tanh_der(self.zs[k][i])) 19 20 for i in range(len(self.ws)): 21 prev = self.act[i] 22 for j in range(len(self.ws[i])): 23 self.grad_b[i][j] = delta[i][j] 24 for k in range(len(prev)): 25 self.grad_w[i][j][k] = delta[i][j] * prev[k]
Ah the backprop. I think this is the hardest part from our code. Assuming that you have read the previous section, let’s interpret the code with the corresponding math equations.
Recall although we do not calcuate the loss per individual parameter, we do
have to calculate the gradients. To start, I decided to put the initial values as \(0\).
To store them, we can do the same just as we randomized the initial weights
and biases because ultimately, each gradient will have a matching parameter.
Continuing, ypred is the prediction, and that is the first element from the
last activation layer. This will vary when we have numerous outputs,
but since this is a simple regression problem and we only have one
output, this is good. delta here is the delta term for our hidden layers.
Let’s start by setting them as \(0\) at first and ypred - y for the last
layer.
Here comes the fun part. In Python, the range function can be used as
range(start, stop,step). Here, we start from the last hidden layer
and we subtract \(2\) due to indices. Then, we will stop when we reach \(-1\) by going
backward by \(1\). We will loop through them for our delta terms. Recall the
formula that we derived below. \[ \delta _i^{[l]} = \frac {\partial L}{\partial a_i^{[l]}} \frac {\partial a_i^{[l]}}{\partial z_i^{[l]}} = g'\left (z_i^{[l]}\right ) \sum _j w_{ji}^{[l]} \delta _j^{[l+1]} \] As the formula suggests, we will add the
weights and the corresponding following delta term with the for loop for \(\frac {\partial L}{\partial a_i^{[l]}}\).
Then, we will mulitply it by the derivative of the actiavtion function, which
is \(\tanh \) for our example. Finally, we can store our delta terms that are
calculuated. Next, we can now use them to calculate the gradients.
\begin{align*} \frac {\partial L}{\partial w_{ij}^{[l]}} &= \delta _i^{[l]} a_j^{[l-1]} \\ \frac {\partial L}{\partial b_i^{[l]}} &= \delta _i^{[l]} \frac {\partial z_i^{[l]}}{\partial b_i^{[l]}} = \delta _i^{[l]} \end{align*}
As the formula suggests, for each parameters, will calculate the gradients again with for the loop. This is it for the backprop!
1def update(self, lr): 2 for i in range(len(self.ws)): 3 for j in range(len(self.ws[i])): 4 for k in range(len(self.ws[i][j])): 5 self.ws[i][j][k] -= lr * self.grad_w[i][j][k] 6 self.bs[i][j] -= lr * self.grad_b[i][j] 7 8def train(self, train_x, train_y, epochs, lr): 9 losses = [] 10 for epoch in range(epochs): 11 total_loss = 0 12 13 for x, y in zip(train_x, train_y): 14 ypred = self.forward(x) 15 total_loss += loss(y, ypred) 16 self.backprop(y) 17 self.update(lr) 18 19 average_loss = total_loss / len(train_x) 20 losses.append(average_loss) 21 22 if (epoch + 1) % 100 == 0: 23 print(f"Epoch: {epoch+1}/{epochs}, Loss: {average_loss}") 24 25 return losses
Now we need to update. This part is similar with our previous example, but for each parameter, we will update it by subracting the product of the corresponding gradient and the learning rate. The training part is again very similar to our previous example. We will begin our loss with \(0\), the we will predict, accumulate the loss, backprop, update, and print every hundred epoch.
1 def show_params(self): 2 for i, (w, b) in enumerate(zip(self.ws, self.bs)): 3 layer = "OUTPUT" if i == len(self.ws) - 1 else f"LAYER {i+1}" 4 print(layer) 5 for j, neuron_w in enumerate(w): 6 weights_str = ", ".join(f"{x:.4f}" for x in neuron_w) 7 print(f" Neuron {j+1}:") 8 print(f" Weights: [{weights_str}]") 9 print(f" Bias = {b[j]:.4f}") 10 print() 11 12mlp = MLP([2, 8, 8, 1])
This part is for printing the parameters, and as you can see, it is pretty straight forward. I ended up not using it because I don’t think it is that effective at this point, but why not just have it. Now we create an instance for the MLP class with two hidden layers with eight neurons each.
1losses = mlp.train(train_x, train_y, epochs=2000, lr=0.01) 2print() 3 4for x, y in zip(train_x[:10], train_y[:10]): 5 print(x, y) 6 print(f" ypred: {mlp.forward(x)}") 7 8plt.plot(losses) 9plt.xlabel("epoch") 10plt.ylabel("loss") 11plt.show() 12 13#mlp.show_params()
We will train with the learning rate of \(0.01\) and \(2000\) epochs. We then print the first ten p9redictions and plot the loss.
1test_loss = 0 2for x, y in zip(test_x, test_y): 3 ypred = mlp.forward(x) 4 test_loss += loss(y, ypred) 5 6for x, y in zip(test_x[:10], test_y[:10]): 7 ypred = mlp.forward(x) 8 print(x, y) 9 print(f" ypred: {ypred}") 10 11print("Test Loss:", test_loss / len(test_x))
This part is for printing the loss and prediciton for testing. That was it for the explanation! Now let’s take a look at the results.
Analyzing the Results
Now that we have looked at the source code, let’s analyze the result to see how our vanilla neural net performs.
1losses = mlp.train(train_x, train_y, epochs=2000, lr=0.01) 2print() 3 4for x, y in zip(train_x[:10], train_y[:10]): 5 print(x, y) 6 print(f" ypred: {mlp.forward(x)}") 7 8plt.plot(losses) 9plt.xlabel("epoch") 10plt.ylabel("loss") 11plt.show() 12 13#mlp.show_params()
Epoch: 100/2000, Loss: 0.0006795237745380497 Epoch: 200/2000, Loss: 0.0002780199905385479 Epoch: 300/2000, Loss: 0.00015780133554953304 Epoch: 400/2000, Loss: 9.996884725589866e-05 Epoch: 500/2000, Loss: 6.544552166502369e-05 Epoch: 600/2000, Loss: 4.373246439583505e-05 Epoch: 700/2000, Loss: 3.032175895325618e-05 Epoch: 800/2000, Loss: 2.2295750007135315e-05 Epoch: 900/2000, Loss: 1.7559747007674843e-05 Epoch: 1000/2000, Loss: 1.4723606896641237e-05 Epoch: 1100/2000, Loss: 1.2951440585334302e-05 Epoch: 1200/2000, Loss: 1.1773900112740833e-05 Epoch: 1300/2000, Loss: 1.0936643742787753e-05 Epoch: 1400/2000, Loss: 1.0302935457567813e-05 Epoch: 1500/2000, Loss: 9.79811628836463e-06 Epoch: 1600/2000, Loss: 9.379957580934361e-06 Epoch: 1700/2000, Loss: 9.023360655331247e-06 Epoch: 1800/2000, Loss: 8.712547584218995e-06 Epoch: 1900/2000, Loss: 8.437042172730164e-06 Epoch: 2000/2000, Loss: 8.189547717119376e-06 [1.0, -1.0] -1.0 ypred: -0.9914314084230468 [0.8, -1.0] -0.8 ypred: -0.8026975789984329 [0.8, -0.9] -0.7200000000000001 ypred: -0.7229659139930072 [0.1, -0.8] -0.08000000000000002 ypred: -0.07901899113208088 [0.3, -0.3] -0.09 ypred: -0.08931568886644314 [-0.7, -0.1] 0.06999999999999999 ypred: 0.07003157242570901 [-0.6, -0.5] 0.3 ypred: 0.29863205953269367 [0.7, 0.8] 0.5599999999999999 ypred: 0.5613061391881997 [-0.5, -0.4] 0.2 ypred: 0.19629911113434118 [0.9, 0.8] 0.7200000000000001 ypred: 0.7218073084017129
![]()
I would say this is pretty good! The loss is constantly decreasing and the prediction seems to be pretty good. Looking at the graph, the model seemed to have little idea at first, but instantly learned to see the holistic view as suggested by the sudden drop.
1test_loss = 0 2for x, y in zip(test_x, test_y): 3 ypred = mlp.forward(x) 4 test_loss += loss(y, ypred) 5 6for x, y in zip(test_x[:10], test_y[:10]): 7 ypred = mlp.forward(x) 8 print(x, y) 9 print(f" ypred: {ypred}") 10 11print("Test Loss:", test_loss / len(test_x))
[-0.9, -0.3] 0.27 ypred: 0.27224246623233583 [0.9, 1.0] 0.9 ypred: 0.889162612694173 [0.1, 0.5] 0.05 ypred: 0.050824340568330406 [1.0, 1.0] 1.0 ypred: 0.9737260427905209 [0.6, -0.8] -0.48 ypred: -0.4811987339447982 [0.6, -0.3] -0.18 ypred: -0.17587974783858806 [0.6, 0.4] 0.24 ypred: 0.23577905969192914 [0.5, -0.8] -0.4 ypred: -0.4007379739663608 [0.5, 0.2] 0.1 ypred: 0.0970149401772119 [0.9, -0.2] -0.18000000000000002 ypred: -0.18071247280108016 Test Loss: 1.1432408362631248e-05
Our test also seems pretty good. The loss seems small because our validation set uses small numbers, but we could see that the prediction somewhat matches the real values!
Now looking back at our model, this is again good for the limited range as suggested by the result above. However, there is a reason for that. Recall that we shuffled the dataset and seletcted \(80\%\) for training and \(20\%\) for testing there. Therefore, for the limited range that the model is trained in, the model did pretty good job. However, when I first tried making dataset by splitting the training and validation as first \(80\%\) and latter \(20\%\), our model did poorly. This is because the model learned the left side well, but it does poorly in predicting the right side since it has no idea what’s going on there with absolutely no training example. However, the good news is that our neural net is scalable unlike the previous example! By increasing the range for datasets, and could see whether the model underfits or overfits.
We now covered two models from scratch! Now that we know what happens in deep neural nets, let’s build our own language models!