Linear regression is the machine learning equivalent of that one friend who always shows up on time, doesn’t over complicate things, and still manages to get the job done brilliantly. It’s one of the simplest forms of regression models out there, but don’t let that fool you it packs a punch in terms of predictive power. Of course, like any good performance, it all depends on how you train it.
Enter hyperparameters, those mysterious dials and knobs you tweak before training begins. In machine learning, hyperparameters aren’t learned by the model itself; they’re set by you, the human in the loop, which means you’re now officially part of the learning process. Congratulations!
Now, if you’re sticking with good old standard linear regression, you’re off the hook—there are no hyperparameters to worry about. But once you bring in the heavy hitters like Ridge Regression, Lasso Regression, or the ever-so-balanced Elastic Net, hyperparameters become mission-critical. These aren’t just settings—they’re the fine-tuning tools that help your model walk the tightrope between accuracy and overfitting.
Before We Tune the Hyperparameters, Let’s Talk Data and the Fit That Makes or Breaks Your Model
Before diving headfirst into the world of hyperparameter tuning, it’s essential to take a step back and look at the real star of the show: your data. After all, even the most sophisticated machine learning models can’t make magic out of messy inputs. So, let’s start with the basics—training and testing datasets.
Typically, we split our dataset into two parts: the training set and the test set. The training set is where your model gets its education it learns patterns, relationships, and hopefully not too many bad habits. The test set, on the other hand, is like a surprise pop quiz. It helps us evaluate how well the model generalizes to unseen data, or as we like to call it, “the real world.”
Now, here’s where things can go hilariously wrong.
Overfitting happens when your model becomes too good at learning the training data so good, in fact, that it starts memorizing the noise instead of learning the patterns. It’s like that one student who can recite the textbook word-for-word but completely blanks when asked to apply the knowledge. Sure, the training performance is stellar, but the moment you feed it new data? Chaos.
Underfitting, on the other hand, is the opposite end of the spectrum. This is when your model is… well, not even trying. It fails to capture the underlying trends in the data and performs poorly even on the training set. If your model can’t pass the open-book exam (a.k.a. the training set), don’t expect it to ace the final. In such cases, the solution usually involves increasing model complexity perhaps adding more features or using a more expressive algorithm.

Understanding overfitting vs. underfitting is crucial in machine learning model evaluation. It’s a balancing act: you want your model to generalize well without being either too rigid or too naive. Only once you’ve nailed that balance does it make sense to reach for the hyperparameter knobs and start tuning like a pro.
Now Tuning: Linear Regression Hyperparameters (a.k.a. Flipping the Right Switches)
In scikit-learn, the go-to Python library for machine learning, the LinearRegression class may seem minimalistic, but it does offer a handful of important knobs worth adjusting especially when you’re aiming for optimized performance or reproducibility.
Here’s what’s under the hood:Here’s what’s under the hood:
sklearn.linear_model.LinearRegression(
fit_intercept=True,
copy_X=True,
n_jobs=None,
positive=False
)
Let’s decode what these hyperparameters actually do without sending your brain into a recursive loop:
fit_intercept (default = True)
- This one asks a simple question: Should we calculate the intercept or not?
- If set to
False, the model assumes the data is already centered (i.e., the mean is zero), and it skips fitting an intercept. Great for when you’ve already done your preprocessing homework. Otherwise, leave it asTruebecause most real-world data isn’t so well-behaved.
copy_X (default = True)
- Should the input feature matrix
Xbe copied before fitting? - If you’re working on large datasets and want to save memory (or live dangerously), setting this to
Falsemay help. But if you’re not sure, it’s safer to let it copy. After all, corrupted data is a machine learning horror story no one wants.
n_jobs (default = None)
- This parameter controls parallel processing.
None= single-threaded (the default behavior).-1= unleash the full power of all your CPU cores.
- It’s only effective when solving sufficiently large problems, such as multi-target regression or when using
positive=Truewith sparse data. In other words, if your dataset is tiny, this switch won’t do much but for heavy lifting, it’s a game changer.
positive (default = False)
- Set this to
Trueif you’re building models where negative coefficients just don’t make sense like predicting physical quantities (e.g., price, mass, time) that can’t be negative. It forces all regression coefficients to stay in the positive zone. But fair warning: this only works with dense arrays. If your data is sparse, this parameter politely bows out.

It’s time to peek under the hood and understand what the model has actually learned. Here are some of the essential attributes that scikit-learn exposes after fitting the model, and why they matter for your analysis and interpretation:
coef_ (array)
This is the heart of your model, the estimated coefficients corresponding to each feature.
- If you’re predicting a single target,
coef_is a simple 1D array with one coefficient per feature. - For multiple targets, it becomes a 2D array where each row corresponds to a target and each column to a feature.
rank_ (int)
- This tells you the rank of the input feature matrix
Xbasically, a measure of how many independent features you truly have (whenXis dense). - If your data is collinear (think: redundant or linearly dependent features), this number will be less than the total number of features, hinting at potential multicollinearity problems.
singular_ (array)
- An array containing the singular values of your feature matrix
X(when dense). - Singular values give you insight into the numerical stability and conditioning of your data matrix important for diagnosing ill-conditioned datasets that might mess with your model’s accuracy.
intercept_ (float or array)
- The intercept term the baseline prediction when all features are zero.
- If you set
fit_intercept=Falseduring model initialization, this will be zero by design
n_features_in_ (int)
- Simply the number of features your model saw during training.
- Handy for sanity checks and ensuring your input data is consistent.
feature_names_in_ (array)
- If your input data
Xhad feature names (all strings), this attribute stores them for easy reference. - Great for keeping track of which features correspond to which coefficients especially useful in datasets with many variables.
While LinearRegression may not have as many knobs as Ridge or Lasso, even these subtle settings can have a significant impact especially when scaling up to production or working with highly sensitive datasets.
Remember, hyperparameters are like seasoning: too much or too little can throw off the entire dish. But when tuned just right? Chef’s kiss.
Here are a few of the usual suspects in the world of linear regression hyperparameter tuning:
- Regularization Strength (Alpha): Think of alpha as the model’s conscience. A higher alpha discourages overly dramatic coefficients, nudging the model toward simplicity and better generalization.
- Learning Rate (for Gradient Descent): This one controls how aggressively your model updates itself during training. Too slow, and you’re watching paint dry. Too fast, and you might miss the sweet spot entirely.
- Normalization or Standardization: Scaling your features ensures that all input variables get a fair shot at influencing the outcome—no favoritism based on magnitude. It’s especially helpful when using gradient descent, where uneven scales can send optimization off the rails.
It’s time to give it a performance boost. Enter GridSearchCV, the data scientist’s equivalent of a personal trainer for machine learning models.
GridSearchCV helps us automatically search through combinations of hyperparameters to find the best possible set, so your model doesn’t just work, it shines.
from sklearn.model_selection import RandomizedSearchCV
import pandas as pd
from sklearn.linear_model import LinearRegression
from sklearn.metrics import r2_score
from sklearn.model_selection import train_test_split
dataset_url = "datasetengine"
# Independent variables
X = df_cars[["Year", "Kilometers_Driven", "Mileage",
"Engine", "Power", "Seats"]]
# Dependent variable
Y = df_cars["Price"]
# Split the data
X_train, X_test, y_train, y_test = train_test_split(
X, Y, test_size=0.2, random_state=0)
model = LinearRegression()
param_space = {'copy_X': [True,False],
'fit_intercept': [True,False],
'n_jobs': [1,5,10,15,None],
'positive': [True,False]}
random_search = RandomizedSearchCV(model, param_space, n_iter=100, cv=5)
random_search.fit(X_train, y_train)
# Parameter which gives the best results
print(f"Best parameters: {random_search.best_params_}")
# Accuracy of the model after using best parameters
print(f"Best Score: {random_search.best_score_}")
Output
Best Hyperparameters: {'positive': False, 'n_jobs': 1, 'fit_intercept': True, 'copy_X': True}
Best Score: 0.7510045389329963
What we have done is technically correct syntactically valid, as we say in the machine learning world but let’s just say it won’t exactly set your model’s performance on fire. Think of it like installing a spoiler on a bicycle: sure, it fits… but you’re not winning any races
If you’re truly chasing performance improvements, especially in real-world regression problems, your best bet lies in using regularized linear models like Ridge Regression, Lasso, or the ever-diplomatic Elastic Net. These models don’t just fit they generalize, which is the whole point.
For good old LinearRegression (the vanilla flavour), you can still tweak a couple of dials like fit_intercept and positive to see minor gains. A simple loop or a small GridSearchCV can do the trick here. Just don’t expect miracles; think of it more as fine-tuning your espresso shot rather than rebuilding the coffee machine.
Up Next: Ridge, Lasso, and Elastic Net – The Regularization Trifecta
We’ve just scratched the surface of what regularization can do. In our next blog, we’ll dive into the differences between Ridge Regression, Lasso, and the hybrid genius known as Elastic Net Regularization. Spoiler: it’s where math meets finesse.

Stay tuned—your models are about to get leaner, smarter, and more generalizable.




One response to “Linear Regression and Hyperparameter Tuning: Complete Guide”
[…] Linear Regression and Hyperparameter Tuning: Complete Guide […]