"No one is harder on a talented person than the person themselves" - Linda Wilkinson ; "Trust your guts and don't follow the herd" ; "Validate direction not destination" ;

March 10, 2017

IoT Training

It would be useful to do a few student projects to understand the fundamentals and hands-on training. This is next learning plan from June. Bookmarking a few interesting labs.
Happy Learning!!!

March 03, 2017

Day #56 - Deep Learning Class #3 Notes - Training Deep Networks

Part I
Parameters Overview
  • Multi Layer Perceptron (Mean Square Error, Weight Decay Term - prevent overfitting, regularization term)
  • Error function non-convex loss function - Gradient Descent
  • Saddle points (Minimum along few dimensions, maximum along other dimensions), They can be problem in Deep networks. Train Deep networks to avoid saddle points
  • Vanishing Gradient Problem - Identify weights, value becomes small updates become slow. Chain rule in backpropagation. Gradient update that reaches earlier layers will be very low
  • Alexnet had only seven layers, Best of network use only 7~8 layers
  • Exploding gradient, terminate the value when it is greater than threshold
  • Mini Batch SD - Batches of GD (20 / 100 points), Average 20 points and update all layers in network. SGD with batch size
  • Iteration - Whenever a weight update is done
  • Epoch - Whenever training set is used once
  • Momentum - During GD, idea is example of blind person navigating mountain range
  • Momentum(Useful - Find the local minima or some other local minima or better minima, Highly Elliptical momentum useful /Not Useful - For shperical contour plot)
  • Spherical - Normal will take you directly to centre
  • Contour Plot - cross section of mountain
  • Nesterov Momentum - One step further in direction of step we need to take, Works very well in practice (Intermin update, compute weights)
Part II
Choosing Activation Function
  • Different Activation functions Sigmoid, tanh, Relu, Leaky Relu, maxout
  • Sigmoid (Between 0 and 1) - By Applying sigmoid this range will be reduced, Zero output will be eliminated by tanH. Brings non-linearity in network
  • TanH - Belong to logistics family of functions (-1 to +1), Used even today
  • Relu - Most popular - Rectified Linear Unit, Actually Linear on +ve Side, Negative side Zero. Max(0, x). For -ve output Relu will make it Zero. Relus is linear on positive axis, Default for images and Videos
  • Leaky Relu - y = x if x > 10, Going to let small amount of x pass thru
  • maxout - Given layer of neurons, groups of 10. Max of batch of neurons is the maximum
  • Softmax - Ensure each activation lies between 0 and 1
  • Hierarchical Softmax - Requires word2vec
  • Relu mostly for images and videos, If too many dead units try leaky Relu
  • RNN, LSTM still sigmoid and tanH is used
Part III
Choosing Loss Function
  • Loss Functions / Cost both mean the same
  • MSE - Gradient is simple
  • Cross Entropy Loss function
  • Entropy - -SumPiLogPi
  • Binary cross Entropy
  • Negative Log Likelihood
  • Softmax for binary same as Sigmoid
  • Start with NLL, Minimise NLL given particular activation function
  • KLDivergence measure distance between two distributions
Part IV
Choosing Learning Rate
  • Convex function GD will always take you to local minimum
  • Always GD reached in one step for correct learning rate
  • Hessian - second derivative matrix (Optimal learning rate - inverse of Hessian)
  • Gradient is a vector not a single value
  • Optimal Learning rate - Eigen values of Hessian
  • Adaptive is best approach to chose learning rate
  • Adagrad is one such method
  • Slow on Steep clif, On Flat Surface long long approach
  • RMSProp - Root Mean Square Prop 
  • Adam - Most popular method today (Default)
  • Momentum + Current Gradient - Adam
  • AdaDelta another similar method
  • To chose - http://sebastianruder.com/optimizing-gradient-descent
  • SGD - Mini batch SGD
  • Choice for Training - SGD + Nestrov momentum, SGD with Adagrad / RMSProp / Adam
Part V
  • Math of Backpropagation
  • Backpropagation uses GD
  • Issues with GD / Training GD
  • Using Learning Rate / Optimization
Part VI - Regularization
  • Training DL is GD
  • ML and Optimization difference is Generalization
  • ML best performance tomorrow (Work well tomorrow, generalization is important)
  • Regularization methods incorporated for Generalization performance
  • Training accuracy increases, Test Accuracy Decreases (Point to stop)
Part VII - When to Stop
  • Train Epoch, Lower learning rate and again Train
  • Maxweight change less than particular row
  • Weight decay term in Error function itself
  • L2 Weight Decay (Add Square of weights)
  • L1 Weight Decay (Absolute Value of Weights), Sparse Solutions
  • Drop Out - In each iteration, In each mini batch, In every layer randomly drop certain % of nodes. This gives excellent regularzation performance
  • Ensemble of different models 9Similarity of Random Forests)
  • DropConnect (extension of Drop out)
  • Add Noise (Data Noise) - Gaussian, Salt and Pepper Noise
  • Batch Normalization Layer (Recommended) - All implemented libraries 
  • Shuffle your inputs
  • Choose mini-batch such that network learns faster
Curriculum Learning
  • Provide slides and figure the course out
  • Lots of data + Lots of computing for Deep Learning Success (Google / FB)
  • Unsupervised Learning is approach by Facebook for Data Analysis
  • Data Programmatically - NIPS Machine Learning Conference
  • Data Augmentation (Change illumation in data, Reduce intensity of pixels, Train Network with all kinds of data - Mirror, Noise, Artificial Images)
Target Values
  • Binary classification problem ? +1 and -1
Weight Initialization
  • GD works and takes you to different local minima
  • Starting defined by how you initialize the network
  • Never Initialize to Zero
  • Recommended ways - Xaviers Initialization
  • For every layer in network get weights randomly from uniform distribution

Happy Learning!!!

February 10, 2017

Day #55 - Markov chains Basics

This post is from my notes. I had bookmarked some interesting answers on understanding Markov chains.

What is a Markov chain?
The simplest example is a drunkard's walk (also called a random walk). The drunk might stumble in any direction but will move only 1 step from the current position.

The ink drop in a glass of water example

Imagine a traffic light with three states: yellow, green, red; however, instead of going Green-> Yellow-> Red at "fixed intervals", it would go at any color at any time.(randomly - Imagine a dice with 3 color and you throw it and decide what color it will be next).   Alternatively, imagine you are in certain color, say green. If you don't allow to be in the same color again, flip a coin. If it is heads go to red, and if tails go to yellow.

So to make a "chain" we just feed tomorrows result back into today. Then we can get a long chain like rain rain rain no rain no rain no rain rain rain rain no rain no rain no rain a pattern will emerge that there will be long "chains" of rain or no rain based on how we setup our "chances" or probabilities.

Markov Chain - Khan Academy
  • Hidden blue prints of nature / objects around us
  • Once you begin each sequence will converge to one ratio
  • First order and second order model defined by Claude Shannon
Happy Learning!!!

February 07, 2017

Day #54 - Fundamental Concepts - Artificial Neural Networks

Referenced Articles - Link

One liner definitions
  • Image - Represented as RGB Matrix with Height and width = 3 color channels X Height X width
  • Color represented in [0,255] Range
  • Kernel - Small Sized matrix consists of real-valued entries
  • Activation Region - Region where features specific to kernel detected in input
  • Convolution - Calculated by taking dot product of corresponding values of kernel and input matrix certain selected coordinates
  • Zero Padding - Systematically adding inputs to adjust size based on requirements
  • Hyperparameter- Properties pertaining to the structure of layers and neurons (spatial arrangement, receptive field values called hyperparameters). Main CNN hyperparameters are R - Receptive Field, Zero Padding - P, input volume dimension ( Width X Height X Depth) and Stride Length (S)
  • Convolutional Layer - Convolution operation with input filters and identifying the activation region. Convolutiuon Layer output - ReLu (Activation Values)
  • ReLu - Rectified Linear Unit Layer. Most commonly deployed activation function for output of CNN neurons. max(0,x)
  • ReLu is not differentiable with origin so we use Softplus function ln(1+e^x). Derivative of Softplus function is sigmoid function
  • Pooling - Placed after convolution. Objective is downsampling (reduce dimensions)
  • Advantages of downsampling
    • Decreased size of input for upcoming layers
    • Works against overfitting
  • Pooling takes sliding window across input transforming into representative values. Transformation performed by taking maximum value in observable window (max pooling)
Happy Learning!!!

February 03, 2017

Day #53 - Tech Talk - Nikhil Garg - Building a Machine Learning Platform at Quora - MLconf SF 2016


Keynotes from Session

Machine Learning Platform - Collection of systems to sustainable increase the business impact of ML at scale

Build or Buy
1. Degree of Integration with the product. Delegation of components
2. Support for Production Systems (cannot outsource business logic to outside platforms)
3. Blurry line between experimentation & production
4. Leverage Open source in an open manner
5. Commercial platforms are not super valuable - Can often train most models in single multi-core machine
6. Blurry line between ML & Product Development (Inhouse tools for monitoring/training / deploying etc..)
7. ML is Quora's core competency

Machine Learning Models Deployed


Machine Learning Use Cases


Happy Analytics!!!

February 02, 2017

Machine Learning Quotes

Quote #1 - "In Markov model our assumption is future state depends on only current state, not any other previous states"

Quote #2 - "In Bayes, we have naive assumption the current term is independent of the previous term - Naive assumption"

Happy Learning!!!

January 21, 2017

Day #52 - Deep Learning Class #1 Notes

AI - Reverse Engineering the brain (Curated Knowledge)
ML - Machine Learning is subset of AI. Teaching Machine to Learn

Deep Learning - Rebirth of Neural Networks
  • Multiple layer of neurons
  • Directed Graph
  • First Layer is input layer
  • Last Layer is output layer
  • Intermediate layer is hidden Layer
  • Deep Learning is inspired by human brain
  • In Deep Learning features are learnt
  • Gradient Descent - Process of making updates in NN
  • Neural Networks is discriminative approach
  • Neurons in neural networks end up in becoming feature selectors
Discriminative Classifiers - Logistics, SVM (uses kernel for non-linear classification), Decision Trees
Generative Model - Naive Bayes

Types of Neural Networks
  • Autoencoders for dimensionality reduction
  • CNN Convolutional NN
  • RNN Recurrent NN
Interesting Deep Learning Demo Sites Discussed
imsitu.org
cloudcv.org

Happy Learning!!!

January 18, 2017

Neural Networks - Learning Resources

Happy Learning!!!

Interesting Data Science Projects

Happy Learning!!!

January 13, 2017

Day #51 - Neural Networks

Happy New Year 2017. This post is on Neural Networks.

Neural Networks
  • ANN - inspired by biological networks, Modelling network based on neurons
  • Key layers - Input Layer, Hidden Layer, Output Layer
  • Neural networks that can learn - Perceptrons, backpropagation networks, Boltzaman machines,recurrent networks
  • In below example for XOR implementation we use backpropagation 

Implementation overview
  • Initialize the edge weights at random (we do not exact weights we chose randomly, By training we find exact values)
  • Calculate the error - we have some training data and some results - Supervised learning, Calculated output not logical output, Error term present 
  • Calculate the changes of edge weights and update the weights (backpropagation process), Calculate edge weight changes and update accordingly
  • Algorithm terminates when error rate is small

Happy Pongal & Happy Learning!!!