Showing posts with label Activation Functions. Show all posts
Showing posts with label Activation Functions. Show all posts

Sunday, June 1, 2025

AI / ML - History of Activations

 In one of my favorite film (Tamil) which deals with the greatness of a 5th Century  spiritual master  (Bodhi Dharma),  there is a wonderful dialogue - Film - Aezhaam Arivu (which means "7th sense")

" We started losing our science when we started ignoring our history and heritage"

***********************************************************************************

 In 1943, "Step" function was used to have the early neural network just to enable "fire" or "don't fire"  in other word - "give output or keep quiet based on the computation made" (no details covered in this post please). You can appreciate this is the most simplistic way of looking at things but the data scientists those days were sincere to mimick their limited understanding of human brain which they observed that only few of the neurons of the brain were activated at any point of time.

Understandably this milestone made their initial neural network to continue the data processing in neural networks - or what is referred as "forward propagation". However, it lacked the ability to support the reverse calculation of the neural network automatically which is critical to optimize the output. Can you believe the AI experts used to manually work on derivatives (calculus) to handle the back propagation since that was not enabled by "step" function ? It was cumbersome but there were no choice for them.

During 1970 - 80, a breakthrough was achieved by using sigmoid function in the basic neural networks (also referred as shallow networks which did not have many layers) which helped  optimization when the concept of gradient descent (a method to do the reverse calculation / backward propagation automatically) was developed. 

However, since - as we know - sigmoid function returns a value 0 to 1, there were challenges of "vanishing gradients" during the backward propagation process so the whole idea of optimization was stuck up. In the year 1980, a much better activation called tanh function started getting used which gives an output in the range -1 to 1 which helped to avoid the earlier challenge substantially. However still when there are situations of large values of inputs or very small values of inputs, optimization still suffered.

After quite a while which also marked the advent of "deep learning era" (supported by deep neural networks (allowing multiple hidden layers and large language models), during 2010, RELU activation function  (Proud to say one of the co-founder was Vinod Nair - though he lives in Canada) . This activation is very simple (output either maximum of the input value or zero which ever is higher) and the computational cost was minimal.  It has almost  taken out the woes of optimization and was a huge shot in the arm of deep learning. Even today, though there are many other sophisticated activation functions available, when in doubt or unsure, developing community goes for ReLU for hidden layer activation as a safe bet. Is it the best one available today ? It still has issue with returning values zeros some times.

During 2011-15, a very smart variation of RELU (returns either input value or .001 of input value) named as "Leaky Relu" was introduced to adjust against the issues faced by zero value of the computed value since now it wont return zero any more but a tiny value instead  and keep the backward propagation going with non zero values ! 

In parallel, we had softmax function introduced in deep neural network based on the need to return a probablistic output for a chosen set of values. We should be clear that this was more out of need than any logical progression in the history described so far.

After 2015, we have Swish, Mish , GELU and so on  which has made smoother activation possible for ultra deep networks ..and this is not an exhaustive list of activation functions.

 Well, my constant companion ChatGPT gave a nice idea to remember the history for people over 40 years of age (when obviously the neurons start dying quite rapidly)

*************************************************************************

 "Some Teachers Run Like Super Geniuses"

S = Step & Sigmoid

T = Tanh

R = ReLU

L = Leaky ReLU

S = SWISH

G = GELU

**************************************************************************

 Another Memory Anchor - 

First they 'stepped' (binary), then made it smooth (Sigmoid), then centered (tanh), then said 'lets forget curves, just 'cut'' (ReLU), fixed 'dying neurons' (Leaky ReLU) and finally started 'smart, curvy activations' (SWISH, GELU)


 

 

 

 

 

 

 

Thursday, May 29, 2025

AI/ML ----> Paradox - "part 2" - Neural Network Activations

So, the moral of today's blog is that we get out of challenges only to get caught into newer issues. :-)  That is how human evolution too happened right ?? 

Did  by mistake the postscript of previous post  is copy pasted here ? When we deal with aspects of  truth which is so overwhelming, is it not natural to repeat, re-emphasis  and reiterate ?  

Paradox and Contradictions is in all aspects of our lives - so it is quite easy to understand few nuances of AI/ML if we relate it with this grand truth. I am continuing the same theme of last post in this post albeit with a different topic.

So lets deal with "Activation" function which is a "basic" concept in neural networks.

This concept has become a familiar one in machine learning era also when ML engineers started experimenting with neural networks as early as 1943 ! Before going down the memory lane let me give a contextual understanding so that switching gears will be smoother.

 The essential difference between the traditional Machine learning models - like Linear regression (which uses linear equation) & logistics regression (which uses sigmoid function)  and neural network is  what is referred as "hidden layers". Neural network uses hidden layers while the traditional ML models were not sophisticated enough to use it.If there is just 1 hidden layer we call it as "shallow" neural network positioned between input layer and output layer. Hidden layer may contain one or many neurons. (well, it could be a very rare situation to have just 1 neuron in a hidden layer but still technically feasible if practically required). 

Essentially what does such hidden layers do ? 

The hidden layers take an input value from previous layer, process it and pass on its output to next layer. As simple as that. 

Processing here essentially means two components - (1) computing operation (which is referred to calculating Z value) and (2) transforming that value (with the help of an activation function)

 The funny thing about activation is that we have multiple choices based on the requirements - the purpose of model, the volume and pattern of data, the complexity of design and also aspects like budget & infrastructure availability. These days the default activation function used is RELU (which is acronym of Rectified Linear Unit) but there are many other smart activation available too.

Now let me zero in on the intent of this post - paradox - isn't it ? When we say there are multiple options today, the paradox is that we don't use sigmoid function and linear function as activation layer in hidden layers.  Both of them are relevant only for output layers and sigmoid was used until a point of time when RELU was made public (Linear function was never attempted to be used in hidden layer since it won't "transform" the output. why to create hidden layers then ?)

I wanted to go deeper into the history of these activation functions but don't want to make this post any longer since I hit the bull's eye already. Do you see the paradox. The original functions that started off machine learning are out-dated and become useless with modern neural networks. In the next post let me explain the history of various activation functions & after that would like to jump in for another sequel to lament about philosophical aspect of this "activation" and what is the purpose of "activation"  in grand scheme of things.

By the way, with my limited exposure to AI/ML, I am convinced that this subject is not for those who have a dis-taste for philosophy. Come on, if we don't see the connection between the most materialistic things and the abstract aspects of life, I am sorry, you are missing life.

 So Lets jump to philosophy after exploring history.. ok ? Stay tuned.