Skip to main content

Command Palette

Search for a command to run...

The ML Script : Week 2

Week 2 of my ML journey is done.

Published
•4 min read•View as Markdown
The ML Script : Week 2

Hello World!

It feels nice to come back after a week, open Hashnode, and write about my experience. This past week was not quite as productive as Week 1 was, but it's fine—who cares? Progress is progress.

This week, I dove into Linear Regression (LR). I started with Simple LR (1 feature + 1 label) and moved into Multiple LR (more than 1 feature + 1 label). Just a reminder: "feature" refers to the input columns and "label" is the output column.

I am still trying to wrap my head around the fact that building the model only takes about 7 lines of code, but the math behind it is so unreal and amazing. I learned some important terms related to LR and explored model evaluation techniques. There isn't much to tell yet, but there is so much more to learn. I’m just hoping my brain can keep up and remember the old stuff!

So, let's get into my top 3 learnings for this week:


1. The Transition to Multiple LR

The transition from Simple to Multiple LR was insane. In Simple LR, the loss function equation is straightforward. In Multiple LR, they don't exactly look like brothers—one is simple, and the other is in matrix form.

After completing Simple LR, I could hardly keep up with the math behind Multiple LR. But after a while, everything started making sense. I even tried to derive the equation for it (though I had to look at my notes a few times as I got off track).

2. Why Standardization?

Standardization means converting the values of integer and float features into the same scale. In simple words, standardization involves calculating the distance between the mean of a feature and the different values in that feature in terms of Standard Deviation (SD).

If the mean is $x$ and the SD is $s$ for an age column, let's say there is an entry of 26. After standardization, the value 26 is replaced by $n \times s$. It represents the distance between the mean age and your current value in terms of SD. You read it as: "Age 26 is n standard deviations away from the mean."

Of course, you know what SD is—it tells you how your data is distributed or spread. Low SD means a tightly packed distribution; high SD means data is spread over a large range.

Example: Let's say I have employee data with features like age, salary, location, and experience, and I want to make a classification model to group people by salary (Low, Medium, High). If I don't standardize the age and salary, there will be a weight imbalance. The numerical values in the salary column are way larger than age. The model might become biased, assuming it only needs to look at the salary column to classify someone, ignoring experience or location. A person with zero experience might earn more in an expensive city than an experienced person in a village, but the model won't give enough "weight" to those other features without standardization.

3. fit vs. transform vs. fit_transform

This was a major point of confusion for me during Week 2. This step happens before model training when we split the data into Training and Testing sets.

Previously, I used to standardize the whole dataset at once. I learned that this is a bad practice because it leads to data leakage. You should split the data first and apply standardization separately.

But why?

  • fit: Finds/learns the mean and SD of the data.

  • transform: Uses that learned mean and SD to apply the z-score formula: $z = \frac{x - \mu}{\sigma}$

  • fit_transform: Does both at once.

To follow best practices, we only fit on the training data. We then use that same mean and SD to transform the test data. This ensures the model doesn't "peek" at the distribution of the test set, keeping the evaluation honest. If I get brand-new data later, I have to scale it using the training data's (Mean & SD) parameters before feeding it to the model , we Don't calculate new Mean or SD for Individual data points

(I’ve heard Pipelines handle this magic automatically to prevent the hassle, but I'm still getting the hang of that!)


What's Next?

That was it for this week. There is still a lot more to cover, and I don't think I fully understand it all myself yet! For Week 3, I will tackle Gradient Descent, learn more about LR, and start some projects.

See you in a week. If you don’t find a post, assume I’m slacking off (or stuck debugging).

Till then, let’s get scripting.