- Hierarchical, Divide and Conquer strategy, Supervised algorithm
- Works on numerical data
- Concepts discussed - Information gain, entropy computation (Shanon entropy)
- Pruning based on chi-square / Shannon entropy
- Convert all string / character into categorical / numerical mappings
- You can also bucketize continuous variables
August 31, 2016
Day #29 - Decision Trees
Labels:
Data Science
August 15, 2016
Day #28 - R - Forecast Library Examples
Following Examples discussed. Library used - R - Forecast Library
Happy Learning!!!
- Moving Average
- Single Exponential Smoothing - Uses single smoothing factor
- Double Exponential Smoothing - Uses two constants and is better at handling trends
- Triple Exponential Smoothing - Smoothing factor, trend, seasonal factors considered
- ARIMA
Labels:
Data Science Tips
August 08, 2016
Applied Machine Learning Notes
Supervised Learning
- Classification (Discrete Labels)
- Regression (Output is continuous, Example - Age, Stock prices)
- Past data + Past Outputs used
- Dimensionality reduction (Data in higher dimensions, Remove dimension without losing lot of information)
- Reducing dimensionality makes it easy for computation (Continuous values)
- Clustering (Discrete labels)
- No Past outputs, Only current data
- All Game Playing is unsupervised
- Learning Policy
- Negative / Positive reward for each step
- Inductive (Learn model, Learn from a function) vs Transductive (Lazy learning ex- Opinion from like minded people)
- Online (Learn from every new incoming tweet) vs Offline (Look past 1 Yeat tweet)
- Generative (Apply Gaussian on Data, Use ML and compute Mean / Variance) vs Discriminative (Two sides of Line)
- Parametric vs Non-Parametric Models
Labels:
Class Notes
July 31, 2016
Fifth Elephant Day #2
Fifth Elephant Day #2 - Part I
Session #1 - Content Marketing
Technical Details
Features as input -> Prediction performed (Independent, stateless)
Reasoning - Sequential, Stateful Exploration
Reasoning Problems - Diagnosis, routes, games, crossing roads
Flavours of Reasoning
{subject, predicate, object}
Session #3 - Continuous online learning
Bird of Feathers Session
Deep Learning now
LSTM (Long Short Term memory)
Interword relationships from corpus (word2vec)
Happy Learning!!!
Session #1 - Content Marketing
- Distribute relevant consistent content. Traditional vs Content Marketing
- Delivering content with speed. Channel proliferation (mobile, computers, tablets)
- Intersection of Brands, Trends, Community Interests (Social media post and metrics)
- Data from social media pages, online aggregators
- Computation of term frequency, inverse document frequency
- Using Solr, Lucene for Indexes
- Cosine Similarity
- Greedy Algorithm
- Prediction vs Reasoning problem
- Prediction Problems Evolution
- At Advanced level Deep Learning, XGBoost, Graphical models
Features as input -> Prediction performed (Independent, stateless)
Reasoning - Sequential, Stateful Exploration
Reasoning Problems - Diagnosis, routes, games, crossing roads
Flavours of Reasoning
- Algorithmic (Search)
- Logical reasoning
- Bayesian probabilistic reasoning
- Markovnian reasoning
{subject, predicate, object}
Session #3 - Continuous online learning
- 70% noise in C2B communication
- 100% noise in B2C communication
- Zipfian
- Apriori - Market Basket Analysis
- XGBoost - Alternative to DL
- Bias - Variance Tradeoff
- Spectral Clustering
- Google Deepmind (Used for Air conditioning)
- Bayesian Probabilistic Learning
- Deep Learning - Build Hierarchy of features (OCR type of problems)
- Traditional Neural Network (Fully Connected, lot of degree of freedom)
- Structural causality (Subsystem appears before, Domain knowledge)
- Temporal causality - This and then that happened
- CNN - learning weights
- Spectral clustering
- PCA (reduce denser to smaller)
- Deep Learning - Hidden layers obtained through coarse grained process
- Neural Networks
- Multiple Layers
- Lots of data
Deep Learning now
- Speech recognition
- Google Deep Models on Phone
- Google street view (House numbers)
- Imagenet
- Captioning images
- Reinforcement learning
- Simple mathematical units combine into complex functions
- X-> input, W-> weights, Non linear function of output
- Multiple hidden layers between input and output
- Training hidden layers is challenge
- Define loss function
- Minimize by moving along gradient
- Move Errors back through the network
- Chain rule conception
- Cafee - Configuration file
- Torch - Describe network in lue
- Theano - Describes computation, writes cuda code, runs and gives results
- Used for images
- Images are organized
- Apply Convolutional filter
- For Deep Learning GPU is important
- Convolution (Have all nice features retain them)
- Pooling (Shrink image)
- Softmax
- Other
LSTM (Long Short Term memory)
Interword relationships from corpus (word2vec)
Happy Learning!!!
Labels:
Conference Notes
July 28, 2016
Fifth Elephant Day #1 Notes - Part II
Sessions # - Link
Talk #3 - Machine Learning in FinTech
Use Cases / Scenarios
Happy Learning!!!
Talk #3 - Machine Learning in FinTech
- Lending Space
- Credit underwriting system
- 2% Credit card usage
- 65% of population < 27 yrs
- Digital foot print (mobile)
- Identity (Aadhar)
Use Cases / Scenarios
- Truth Score (Validity of address / person / sources)
- Need Score (Urgency / Time to respond application)
- Saver Score (cash flow real-time analytics)
- Credit Score (Debt to income)
- Credit awareness score
- Continuous risk assessments
- For Safety driving using smartphone sensors
- Spatial / location data
- Road traffic injuries due to distracted driving
- Phone usage - 4x crash risk
- Speedy driving - 45% car crash history
- Driving behavior analysis / driving feedback
- GPS + Inertial Navigational sensors (Accelerometer / Gyroscope / Magnetometer)
- Drive detection
- Event detection
- Collision detection
- Drive summarization and scoring
- Risk modelling
- Events, location of events, duration of events
- Sensors
- Availability - wide variety across devices
- Raw Data - noisy, unevenly spaced time series
- Events - Time scales, combination of sensors
- Model building - Labelled vs unlabelled data, feature engineering
- Algorithms - Stream / batch efficiency
- Cluster data
- Eliminated uninteresting time periods
- Classification / Regression models
- Spectral clustering
- Crop rotation literacy
- Data curation, Query tools on data product
- Visualization and plotting of Agricultural data
- Using Image comparison for Big Cat Counting
- Predicting Big Cat Areas (Territories)
- Observe Nature, Frame Hypothesis, Design Experiments
- Confront with competing hypothesis
- Spacegap program
- Markov chain Monte-Carlo technique
Labels:
Conference Notes
Fifth Elephant Day #1 Notes - Part I
Sessions # - Link
Talk #1 - Data for Genomic Analysis
Great talk by Ramesh. I had attended his session / technical discussion earlier. This session provided insights on genome / discrepancies in genome sequence leading to rare diseases.
Genome - 3 Billion X 2 Characters
Character variables varies from person to person
Stats (1/10th of probability of cancer)
Baseline risk for breast cancer (1/8),(1/70) ovarian cancer
BRCA1 mutation (5-6 fold increase in breast cancer, 27 fold increase for ovarian cancer)
In India
Talk #2 - Alternative to Wall Street Data
This session gave me some new strategies to collect / analyze data
How to Identify occupancy rate at hotel ?
Data Sources
Lot of opportunity
What is the generative value
Happy Learning!!!
Talk #1 - Data for Genomic Analysis
Great talk by Ramesh. I had attended his session / technical discussion earlier. This session provided insights on genome / discrepancies in genome sequence leading to rare diseases.
Genome - 3 Billion X 2 Characters
Character variables varies from person to person
Stats (1/10th of probability of cancer)
Baseline risk for breast cancer (1/8),(1/70) ovarian cancer
BRCA1 mutation (5-6 fold increase in breast cancer, 27 fold increase for ovarian cancer)
In India
- 35% inherited risk mutation
- 1/25 Thalassemia
- 1 in 400-900 Retinitis Pigmentosa
- 1 in 500, Hypertrophic Cardiomyopathy
- 1 Billion reads - 100GB data per person
- Very similar sequence yet one character might differ
- But reference is 3 Billion long
- Need fast indexing
- Suffix Trees and variations
- Hash table based approaches
- Volume of data
- Funnel down of variety of dimensions
- Triplet Code (Molecule)
- Variants of Triplets nailed down to difference of gnome
- GPU processing / reduce computation time
- Hypothesis Testing
- Stats Models
- GPU Processing to reduce computation time
Talk #2 - Alternative to Wall Street Data
This session gave me some new strategies to collect / analyze data
How to Identify occupancy rate at hotel ?
- Count of cars from parking lots
- Number of rooms lights on
- Take pics of rooms from corner of street and predict based on images collected
- Unconventional ways to think of data collection (Beating the wall street model)
- Checking websites
Data Sources
- Direct data gathering
- Web harvesting
- Primary research
- Look at notice patterns in front of you
- Difference in invoice numbers
- Serial number changes, difference values
Lot of opportunity
- Analyze international markets (India / China)
- COGS
- SG
- ETC
- Scarcity - How widely used
- Granularity - Time / aggregation level
- Structured
- Coverage
- Revenue Surprise Estimates
- Dataset insight / Analysis
- Operating GAAP measures
- Generate money in automated system
- Stock sensitivity to revenue surprises
- Identify underlying ground truth
Happy Learning!!!
Labels:
Conference Notes
July 24, 2016
June 17, 2016
June 15, 2016
Day #26 - R - Moving Weighted Average
Example code based on two day workshop on Azure ML module. Simple example storing and accessing data from Azure workspace
Happy Learning!!!
Happy Learning!!!
Labels:
Data Science Tips
June 01, 2016
Day #25 - Data Transformations in R
This post is on performing Data Transformations in R. This would be part of feature modelling. Advanced PCA will be done during later stages
Data Normalization in Python
Happy Learning!!!
Data Normalization in Python
Happy Learning!!!
Labels:
Data Science Tips
Subscribe to:
Posts (Atom)



























