"No one is harder on a talented person than the person themselves" - Linda Wilkinson ; "Trust your guts and don't follow the herd" ; "Validate direction not destination" ;

November 28, 2021

Everything is not same - Perspectives and Clarity matters

20 years of __________________________________________

  • 20 years of experience = 20 years of the same project / different projects?
  • 20 years of experience = 20 years of same role / multiple roles?
  • 20 years of experience = 20 years of services/ product building
  • 20 years of experience = 20 years of 9-6 or 9-12 ?
  • 20 years of experience = How many endless weekends / production go lives
  • 20 years of experience = How many learning migration on skills / domain / data
  • Titles vs experience vs Expertise vs Being aware of true self matters
With experience

  • Balance both journey and current tasks
  • Code to convince someone this is what I meant
  • Code to unblock/find next steps
  • Code to validate this idea works
  • Prototype to share this is feasible
Young folks need time to trust. More than experience connecting with them with all skills/code/experience matters. 

Keep Thinking!!

November 10, 2021

Zillow Machine Learning Fallout

Good read - Link

Machine learning is no silver bullet if you do not consider domain, data, changing environmental factors. A classic case of missing domain knowledge is flagged in this story.

  • Zillow does Real estate - selling, buying, renting, and financing
  • Zillow home value estimation models failed.
  • Assumption - assumption that housing prices would continue to climb without interruption at a stable rate
  • The domain experts warned of issues with the predictions.
  • The business went ahead anyway. Finally, it bombed

Lessons

  • Domain expert warnings considered as Go / No-go for production, not just model accuracy
  • Learn / Incorporate Data Changes to understand changing trends
  • Performing A/B Experiments to understand customer behaviors and leverage optimal values based on outcomes
  • Better model/feature management / keep improving on features / incorporate external factors based on domain expert perspectives #machinelearning #technology #datascience #domainknowledge

Another good read Zillow, Prophet, Time Series, & Prices


WHY IS INTERMEDIATING HOUSES SO DIFFICULT? EVIDENCE FROM IBUYERS

  • Predict that households’ wiliness to pay for liquidity is highest in those markets
  • Sophisticated algorithmic pricing

My Perspectives
  • I love the housing.com approach to rank an area based on amenities, wellness, connectivity
  • Plus a pricing range based on amenities and facilities provided
  • Plus growth potential / Availability
  • Demand vs Supply
A combination of this would suggest a recommended price that a domain expert could adjust based on other external factors. ML is a guideline, not a blind predictor

Keep Thinking!!!

November 09, 2021

Leaf Classification

Leaf Classification

Paper #1 - Plant identification using deep neural networks via optimization of transfer learning parameters

Key Notes

  • 1.2 million labeled images of 1,000 different categories from the ImageNet = one thousand two hundred per class
  • LifeCLEF 2015 - 91,758 labeled images of different plant organs (e.g. flowers, fruits, leaves, and stems), from 1,000 - 91 per class

Parts of Plant

  • Branch 
  • Entire 
  • Flower 
  • Fruit 
  • Leaf 
  • LeafScan 
  • Stem 
  • Overall




  • Increasing the batch size from 20 to 60 improves the overall accuracy
  • 80 patches for data augmentation

Paper #2 - Multi-Organ Plant Classification Based on Convolutional and Recurrent Neural Networks

Key Notes

  • Feature engineering approaches such as Scale-invariant
  • feature transform (SIFT), Bag of Word (Bow), Speeded-Up
  • Robust Features (SURF), Gabor, Local Binary Pattern (LBP).
  • Most generally used features to distinguish leaves of different species
  • Hybrid generic-organ convolutional neural network, abbreviated HGO-CNN
  • Three different sizes: 256, 384 and 512
  • Crop 256 × 256 center pixels
  • Multi-Scale Plant Images Generation
  • During network training, 224 × 224 pixels are randomly cropped from the rescaled images and fed into the network


Keep Exploring!!!

November 04, 2021

Face Swapping - Research Reads

Paper #1 - FaceShifter: Towards High Fidelity And Occlusion Aware Face Swapping

Key Notes

  • Early replacement-based works simply replace the pixels of inner face region
  • GAN-based works  have illustrated impressive results
  • GAN-based network, named Adaptive Embedding Integration Network (AEI-Net)
  • Adaptive Embedding Integration Network (AEINet) to generate a high fidelity face swapping result


  • DeepFakes, and FSGAN all follow the strategy that first synthesizing the inner face region then blending it into the target face

Paper #2 - Face Swapping: Automatically Replacing Faces in Photographs



Paper #3 - Face Detection, Extraction, and Swapping on Mobile Devices

The Face Swap algorithm consists of five main steps:

  • Viola-Jones face detection using Haar-like features [1], Active Shape Model fitting [4], face rotation, skin-tone matching, and smoothing using Laplacian Pyramids [2]. The Viola-Jones face detection uses an OpenCV library [5] to detect faces from a frontal view. 
  • Laplacian Pyramid for face 1
  • Laplacian Pyramid for face 2
  • Laplacian Pyramid after Swapping
  • Final Collapsed Pyramid
  • Image blending Example
  • faceswap-GAN
  • FaceSwap
  • Faceswap Dev
  • Deepfake Faceswap
  • DeepFake Tools

More Reads

Keep Exploring!!!

October 31, 2021

Remembering Facts vs Evaluating Ideas

I find it hard to remember configuration parameters, default settings, metrics. These are key to many certifications. Often we focus on the problem at hand, not specific functions or code to check.

Every definition is custom to each cloud provider and the set of theoretical FAQ questions, syntax specific to language. We neither measure problem solving or domain knowledge but rely on syntax and remembering facts. This is a stark difference between product vs service companies. 

Certification does not necessarily mean you have the skills to build a solution. They merely imply familiarity with a tool/infra. As long as you map your current skills to new skills find the gaps and address you can build the required solution.

Learning is a collection of observations, experiments, experiences, applying your relevant past lessons. It is a compound effect. Building a solution is easy, but thinking from a futuristic perspective marks the difference between a newbie and an experienced techie.

20 years of experience is not working on the same project. The wider you explore bigger the perspective. The more you fail, the more you are aware of different domains/roles. In the end, let it be a collective memory of different experiences. Win or lose enjoy the journey.

I keep coding my logic with a mix of syntax I recollect across SQL, C, Python, R, C#. First, pseudo logic comes to mind. Later the logic is corrected based on StackOverflow answers. Every language has its own way of defining constructs and separators. Am I a bad programmer, mmm maybe... Always there is more to learn :)

Anyways value addition needs to be quantified so you need to pass this too :)


October 30, 2021

AI in Finance

Paper #1 - AI in Finance: Challenges, Techniques and Opportunities

Key Notes

  • Key Areas are capital markets, trading, banking, insurance, leading/loan, investment, asset/wealth management, risk management, marketing, compliance and regulation, payment, contracting, auditing, accounting, financial infrastructure, blockchain, financial operations, financial services, financial security, and financial ethics
  • Classic techniques including logic, planning, knowledge representation, statistical modeling, mathematical modeling, optimization, autonomous systems, multiagent systems, expert systems
  • Modern techniques such as recent advances in representation learning, machine learning, optimization, data analytics, data mining and knowledge discovery, computational intelligence, event analysis, behavior informatics, social media/network analysis
  • Specific business problems, such as market trend forecasting, stock price prediction, credit scoring, fraud detection, financial report analysis, pricing and hedging, marketing, consumer behavior analysis, algorithmic trading, social commerce, and Internet finance.
  • Portfolio planning and optimization: including designing, planning, optimizing and recommending investment portfolios and strategies in a market
  • Forecasting and prediction: including the regression, classification, estimation and prediction of trend (up or down), movement (direction and scale, etc.), value (e.g., price or volatility)
  • Business profiling: including describing, segmenting, characterizing and classifying markets, products, customers, and services.
  • Sentiment and intention modeling: including characterizing, representing, modeling, analyzing and evaluating the polarity, diversity, propensity and their dynamics of customer sentiment and intention 
  • Anomaly detection: such as characterizing, quantifying, detecting, classifying and predicting abnormal, exceptional and changing behaviors, products, patterns, performance





Paper #2 - Enhancing Financial Inclusion using Mobile Phone Data and Social Network Analytics

Key Notes

  • Datasets - call-detail records, credit and debit account information of customers is used to create scorecards for credit card applicants
  • Call-detail records are used to build call networks and advanced social network analytics techniques are applied to propagate influence from prior defaulters throughout the network to produce influence scores
  • predictive model for a target measure of interest (e.g., churn, fraud, default) 
  • sociodemographic  information, such as age, marital status and postcode; debit account activity, including timing and amount of payments; and credit card activity
  • sociodemographic features such as age, marital status and residency as reported at the time of the credit card application are extracted.





Paper - P2P LOAN ACCEPTANCE AND DEFAULT PREDICTION WITH ARTIFICIAL INTELLIGENCE

Key Notes

Features for the first phase are: 

  • debt to Income ratio (of the applicant); 
  • employment length (of the applicant); 
  • loan amount (of the loan currently requested); 
  • purpose for which the loan is taken
  • loan amount (of the loan currently requested); 
  • term (of the loan currently requested); 
  • instalment (of the loan currently requested); 
  • employment length (of the applicant);
  • home ownership (of the applicant. Rented, owned or owned with a mortgage on the property); 
  • verification status of the income or income source (of the applicant. If this was verified by the Lending Club); 
  • purpose for which the loan is taken; 
  • Debt to Income ratio (of the applicant); 
  • earliest credit line in the record (of the applicant); 
  • number of open credit lines (in applicant’s credit file); 
  • number of derogatory public records (of the applicant);
  • revolving line utilisation rate (the amount of credit the borrower is using relative to all available revolving credit);
  • total number of credit lines (in applicant’s credit file); 
  • number of mortgage credit lines (in applicant’s credit file); 
  • number of bankruptcies (in the applicant’s public record); 
  • logarithm of the applicant’s annual income (the logarithm was taken for scaling purposes); 
  • FICO score (of the applicant); 
  • logarithm of total credit revolving balance (of the applicant).

Paper #3 - Behavior Revealed in Mobile Phone Usage Predicts Credit Repayment

Key Notes

  • Mobile phone transaction history prior to the extension of credit, and whether the credit was repaid on time
  • Transition to a postpaid plan
  • Call and SMS metadata

Paper #4 - Data Science in Economics

Key Notes







More Reads - 

Keep Exploring!!!

Solving the right problem at the Right time matters

2013 I was part of the Team that worked on Traffic Forecasting for Retail Stores

  • Multiple stores across geographies
  • Multiple DB’s for each local store

The forecasting system used to run at Enterprise, Synchronize data to local stores with their own internal synchronization jobs. 

  • These jobs were configured to run according to time zones of stores
  • The algorithms were mostly around a weighted moving average, trend + moving average 
  • The forecast job runs leveraging previous data and projects forecasts by the hour for next day, hourly basis patterns
  • The actuals are captured the following day and measured against it
  • In case of data not present sister stores (similar stores) data was leveraged for calculation

Whatever we say as of today measure model drift, missing data features, work at scale, coexist along with existing transaction system was built as server components, custom-built. 

What we missed are

  • Instead of Traffic forecast if we had done a sales forecast it would have helped to apply solutions for both eCommerce and retail giant
  • We had inherent details of out of stock, replenishment alerts. The same could have been used for out of stock forecast per zone, replenishment forecast per zone
  • These real-time reports from RFID could have served as effective forecast opportunities on the same

Sometimes we may have the right technology and architecture but not the right use cases. Now I see the same things ML attempts to do with #kubeflow, #pipelines, #scale but the same problem which was solved with models available at that point in time would take a different set of skills to solve today 😊

Keep Exploring!!!

AI - Education Opportunities

Paper #1 - Strengthening e-Education in India using Machine Learning

Key Notes

Applying different data mining algorithms on the data of the person and suggesting which course is appropriate for him based on his background knowledge


Paper #2 - Personalized Education in the AI Era: What to Expect Next?

Key Notes




Content summarization and question generation Multi-modal content understanding: Human-in-the-loop content design





More Reads

Teaching Machine Learning in K–12 Computing Education: Potential and Pitfalls

Estimating returns to special education: combining machine learning and text analysis to address confounding

Keep Exploring!!!

October 25, 2021

Merlion - open-source machine learning library for time series - Forecasting

Paper - Merlion: A Machine Learning Library for Time Series

Key Notes

  • From Salesforce another forecasting library
  • Merlin includes classic statistical methods, tree ensembles, and deep learning methods. 
  • Merlion implements many diverse models for both forecasting and anomaly detection

Forecasting Algos List

Univariate time series forecasting

  • ARIMA (AutoRegressive Integrated Moving Average)
  • SARIMA (Seasonal ARIMA)
  • ETS (Error, Trend, Seasonality)
  • Prophet
  • Deep autoregressive LSTM

Multivariate forecasting models

  • autoregression algorithm
  • Vector Autoregression






Examples

Documentation

Orbit: A Python Package for Bayesian Forecasting

Orbit: Probabilistic Forecast with Exponential Smoothing

darts is a Python library for easy manipulation and forecasting of time series

Time Series Made Easy in Python

Keep Exploring!!!

October 24, 2021

Indian Startup #Greyorange #AI #DataScience #Robotics #WarehouseAutomation

Indian Startup #Greyorange #AI #DataScience #Robotics #WarehouseAutomation

Useful links for further review

Keep Exploring!!!