,

How Logistic Regression Works: Beginner to Pro Guide

·

Layered data panels with a neural network and rising trend chart

Logistic regression is a supervised machine learning model which is built on the linear regression formula that which is why it is called a regression model. This model is mainly used for classification tasks where the goal is to predict the probability that an instance belongs to a given class or not. Logistic Regression is a statistical algorithm which analyses the relationship between two data factors

Types of Logistic Regression

Binomial Regression

  • Logistic regression is used for binary classification where we use sigmoid function that takes the input as an independent variable and produces a probability value between 0 and 1
  • Where we use sigmoid fuction as the activiation

Multinomial Regression

  • In multinomial regression there can be 3 or more possible unordered types of dependent variable such a ‘bike’, ‘sedan’, ‘SUV’, etc
  • Where we use the softmax fuction as the activation

Ordinal Regression

  • In an ordinal logistic regression there can be 3 or more possible ordered types of dependent variable such as “Low”, “Medium”, or “High”.
  • Here to we use the softmax function as the activation

But before you put your data through, logistic regression comes with a few ground rules:

  1. Independent Observations: Each data point should stand on its own. Think of them like coworkers at a virtual meeting, no side chats or shared secrets.
  2. Binary Dependent Variable: The outcome should be binary (0 or 1). If your target variable has more than two classes (say, cat/dog/hamster), it’s time to invite the softmax function to the party.
  3. Linearity of Log Odds: There must be a linear relationship between the independent variables and the log-odds of the dependent variable. In plain English, the predictors should line up nicely when mapped to the odds of your outcome, not too wild, not too weird.
  4. No Outliers, Please: Logistic regression is a bit of a neat freak. Outliers can throw off the model’s mojo, so make sure your dataset is clean and well-behaved.
  5. Go Big or Go Home: A large sample size is key. With more data, the model learns better and performs more reliably. Small datasets can lead to shaky predictions and models that get cold feet.

If linear regression is the friendly neighborhood line of best fit, then logistic regression is its more decisive cousin who only deals in black and white or rather, zeros and ones. At the heart of logistic regression lies the sigmoid activation function, a smooth, S-shaped curve that takes your wild, wandering predictions and gently nudges them into binary clarity.

Unlike linear regression, which aims to predict continuous values by fitting a straight line through your data, binomial logistic regression is in the business of classification. It’s here to answer yes/no questions like: “Will this email be spam?” or “Is this customer likely to churn?” Instead of fitting a line, we’re fitting probabilities bounded between 0 and 1. You could say it’s regression with commitment issues it never promises more than 1.

consider a medical issue Suppose there is a person, with a blood sugar level of 195, and you do not know whether that person has diabetes or not. What would you do then? Would you classify him/her as a diabetic or as a non-diabetic?

It features a hard threshold decision boundary at a blood sugar level of 200, with clear red and blue markers and a step function line.

based on the boundary, you may be tempted to declare this person a diabetic, but can you really do that? This person’s sugar level (195 mg/dL) is very close to the threshold (200 mg/dL), below which people are declared as non-diabetic. It is, therefore, quite possible that this person was just a non-diabetic with a slightly high blood sugar level. After all, the data does have people with slightly high sugar levels (220 mg/dL), who are not diabetics. using a simple boundary decision method would not work in this case.

Sigmoid activation fuction

this a mathamticall fuction which maps the predicted value to a probabilities

this fuction maps any real value into another value between 0 & 1 this forms a S curve

in the logistic regression we use the concept of the threshold value whihc defins the probability of either 0 or 1 and value that

Plot with a sigmoid curve, which smoothly models the probability of diabetes based on blood sugar level. The curve is centered at 200, showing a gradual transition from 0 to 1 instead of a sharp threshold.

the logistic regression considers the probility of the output. now considering the probility the equation for logistic regression:

P(x) = \frac{1}{1 + e^{-(\beta_0 + \beta_1 x)}}

Where:

  • P(x) : Probability of the outcome (e.g., having diabetes) given the input x (e.g., blood sugar level)
  • \beta_0 : Intercept term
  • \beta_1 ​ : Coefficient (slope) for the predictor x
  • e : Base of the natural logarithm

the best fitting combination of \beta_0 & \beta_1 will be the one which maxixmizes the product mentioend below

(1 - P_1)(1 - P_2)(1 - P_3)(1 - P_4)(1 - P_6) \cdot P_5 P_7 P_8 P_9 P_{10}

A sigmoid curve showing diabetes probability based on blood sugar level.

the equation is called the likeliy hood function

\left[ \prod_{i \in \text{non-diabetics}} (1 - P_i) \right] \cdot \left[ \prod_{i \in \text{diabetics}} P_i \right]

by trying different values of \beta_0  and \beta_1 , you can manipulate the shape of the sigmoid curve. At some combination of \beta_0  and \beta_1 , the ‘likelihood’ will be maximised.

 find the optimal values of \beta_0 and \beta_1 such that the likelihood function is maximized? The optimisation methods used to do that are maximum likelihood estimation, or MLE


Found this useful? Share it:

One response to “How Logistic Regression Works: Beginner to Pro Guide”

Leave a Reply

Keep reading

Discover more from AI and Engineering Insights

Subscribe now to keep reading and get access to the full archive.

Continue reading