"No one is harder on a talented person than the person themselves" - Linda Wilkinson ; "Trust your guts and don't follow the herd" ; "Validate direction not destination" ;

March 29, 2020

Corona Stats - As of March28th

Data Source - Link (As of March28th data)
Case Stats and Growth Trend
Start Date - 2019-12-31
  • Day 69 - 102133
  • Day 81 - 213258
  • Day 84 - 305270
  • Day 87 - 417061
  • Day 89 - 528019
Summary
  • 1st 100K - 69 Days
  • 2nd 100K - 12 days
  • 3rd 100K - 3 days
  • 4th 100K - 3 days
  • 5th 100K - 2 days

Death Stats and Trend
  • Day 1 - 2019-12-31 
  • Day 76 - 5407
  • Day 83 - 11251
  • Day 86 - 16365
  • Day 88 - 20991
Summary
  • First 5K - 76 Days
  • Second 5K - 7 Days
  • Third 5K - 3 Days
  • Fourth 5K - 2 Days

Case Distribution by Country


Fatality



I hope we get through this challenge and recover soon. With global lockdown measures hope we observe downwards trend in the coming weeks.

Good Read - Response to COVID-19 in Taiwan Big Data Analytics, New Technology, and Proactive Testing

Key Points
  • Specific approaches for case identification, containment, and resource allocation
Databases Leveraged
  • Immigration and customs database for travel to Risk Areas
  • Health insurance database for proactively seeking out patients with severe respiratory symptoms 
Risk Categorization
  • Low risk (no travel to level 3 alert areas) 
  • Higher risk (recent travel to level 3 alert areas) 
Inference
  • Real-time alerts during a clinical visit based on travel history and clinical symptoms to aid case identification
  • Tracked through their mobile phone from Self Quarantine

Key Summary Points (Implementation)
  • Avoid partial solutions
  • Learning is critical
Key Summary Points (Lessons Learnt)
  • Extensive testing 
  • Proactive tracing
  • Home diagnosis
  • Monitor and protect health care and other essential workers
Corona Perspectives (July 5th 2020)

Covid Cycle
Unlock Cycle
IT Impact
Carefully we need to plan, bride the gap to address the gaps in the economy, unorganized sectors, poor performing domains. Hope the new normal provide more innovation and newer job opportunities

From NPTEL Lecture Link







Webinar 2 - Link






Keep thinking!!! 

March 21, 2020

Corona impact in Retail

Essentials, Medical and Food supplies, and eCommerce will have a spiked up demand. Clothing / Fashion / Toys / Luxury brands /Smartphone and non-essentials will have an impact leading to reduced sales / temporary closure of stores.

Business Impact
  • Reduced Store Traffic
  • Revisit on Sales Forecasts
  • Temporary Closure of poorly performing stores
  • Supply chain / Manufacturing Delays / Reduced Demands
Alternatives
  • Omni Channel Support
  • Contactless delivery
  • Equip Store associates with Sufficient safety procedures
  • More Sanitizing efforts for store associates/customers
  • A shift for e-commerce mode
  • Use offline data for online personalization
  • Stock up / Align towards products in demand (Healthcare / Medical / Essentials etc)
It will take time to recover and reset/fix the entire supply chain, manufacturing, overcomes direct/indirect job loss, economic impact. Hoping it will be handled well and things will come back to normalcy soon.

These Struggling Retailers May Suffer Their Final Blow From The Coronavirus Lockdown
Challenges
  • Rent and other fixed costs of running
  • Employee Salaries
  • Cash on their balance sheet 
Be Positive!!!
Keep Thinking!!!
Practice Social Distancing!!

Retail AI Landscape

Survey of Retail Landscape, Use cases, Startups




Keep Thinking!!!

March 20, 2020

Day #334 - Lessons Learnt in evaluating SQL 2016 Performance Features

Sharing my lessons on proposing SQL In-Memory table implementation for the product I worked with. I worked with Sunil Agarwal from the SQL product team to evaluate the features, benefits, migration approach, etc.



Happy Learning!!!

SQL 2019 - Interesting Features

SQL 2019 - Interesting Features (Link)

I have a lot of bias for SQL Server. Some SQL 2019 features are awesome. The things I liked are
Query heterogeneous databases with Polybase (Polybase feature was there in 2016 too but the databases supported is not as many as I see now)
  • Polybase provides in SQL Server 2019 through a concept called an EXTERNAL TABLE.
  • External tables are just like SQL Server tables except SQL Server only stores the metadata of the table definition
  • Polybase uses ODBC drivers to connect to sources such as Oracle, Teradata, MongoDB, and SQL Server.
  • Support for SQL, NoSQL
  • Support for ML Engine
  • Support for HDFS
These are promising features. Obviously, there will be some product limitations in the early stages.
Highlights
  • Support for unstructured data
  • Heterogenous database support
  • Schema on read is achieved with external tables
  • SQL wrapper to query both different databases / unstructured data
  • Integration with HDFS
  • ML APIs / Visualization features
Very good move to accommodate / position SQL as an Integration Database engine for heterogenous/unstructured/structured data

Happy Learning!!!

Interesting Product - Intelligent shopping cart

Another interesting product - an intelligent shopping cart

Key features are
  • instore navigation
  • store promotions
  • product suggestions
  • scans and weighs products
  • displays a running tally of purchases 
  • pay on the spot with the cart
Technical Implementation - Link
  • Product Scan (Barcode / RFID)
  • pay/card swipe attached
  • UI to display items scanned / list products
AI Solutions
  • Object Detection 
  • Weight + Object detection for counting
Cons / Concerns
  • Cost of the cart/maintenance 
  • Accuracy of items detected 
  • Accuracy of the count of items for smaller products 
Keep Thinking!!!

March 19, 2020

Day #333 - Deep Learning Guidelines

CI / CD, DL frameworks, Buy vs Develop are different sets of challenges. The more you learn, the more you feel you have a lot to learn :). Learning / doing/debugging/testing everything is part of learning. Keep going!!!

Different levels of learning are required for a different set of challenges.
  • Mastering Keras vs Pytorch vs Tensorflow 
  • Knowing Advanced features of Data Pipelines / Porting in Edge Devices
  • Building end to end the flow of Edge Analytics -> Data Consolidation -> Reporting
  • Deployment of this overall end to end solution
  • Accuracy / Understanding real-world challenges and next incremental  steps
This link provides a good guideline 

The ML tools landscape is very useful



Key Notes
Step #1 - Data
  • Data Storage
  • Data ETL Process (Workflow / Async Process)
  • Data Labelling (Raw Data -> Modelled)
  • Data Versioning

Step #2 - Development / Traning
  • DL Frameworks
  • Source code management
  • Store & Retrieve Results
  • Distributed Training

Step #3 - Deployment
  • Build Tools
  • Web Deployments
  • Monitoring predictions
  • Edge Devices / Custom Hardware Deployment
DL Frameworks


Key Notes
  • Caffe - C++ based (Fintech used Caffe)
  • Tensorflow - Google (Mobile, JS, Scalable Deployment) - Abstraction - Computational Graph
  • Keras - Wrapper on Tensorflow
  • PyTorch - FB product
ML Code Management for Training / Deployment / Serving







Key Lessons
  • Training System (Model Development)
  • Production System (Ready to use Model, Setup)
  • Serving System (Web App or anything that serves model)
On all these three levels there is a certain set of tests run to validate every layer - Train / Model / Production Serving Tests

Infrastructure (Buy vs Build)




Deep Learning Optimization






Data Versioning



Key Lessons
  • Unversioned Data (file system) (L0)
  • Version with a snapshot - Daily data (L1), Data backup with Date
  • A mix of assets and code (L2), JSON or any other labeled storage 
  • L3 - Specialized solution - DVC, Pachyderm, Quill  
Training Neural Nets: a Hacker’s Perspective
Common Coding Mistakes
  • The incorrect shape of tensors
  • Preprocessing inputs incorrectly
  • Incorrect loss function
  • Numerical computation errors (NaN)
Troubleshooting Deep Neural Networks
Troubleshooting Deep Neural Networks

Happy Learning!!!

Distributed Systems - Session #3 - Aurora

Sometimes I felt not connected to the session. Needs a lot of focus and patience to stay connected and focused :)


Key Summary points
  • Amazon early offering EC2
  • Rented out VMs to customers
  • VMM (Virtual Machine Monitors) that run/manage EC2 instances
  • EC2 good for stateless web servers
  • S3 - Scheme for storing large chunks of data (Periodic Snapshots)
  • Disks for EC2 instances - Fault Tolerance (EBS)
  • EBS (Elastic Block Store) - Looks for EC2 instances as it is a harddrive
  • Databases on EBS sends a large volume of data over the network
  • Amount of writes on Network Storage System
  • CPU / Disk space consumption
  • EC2 / EBS are in same availability zone
  • Transaction & Crash Recovery
  • Transaction (Sequence of operations / commands / atomic / ex- bank transfer money between accounts)
  • Reads page from disk
  • Make Changes in local cache
  • Then write changes to disk
  • Log entries describe the transaction
  • Three log records - Modify Operation, Old Value, New Value
  • Aurora is based on MySQL
  • RDS (Database replicated in multiple availability zones)
  • All the transactions mirrored to other databases (EBS Servers)
  • Multiple copies managed and updated to keep everything in sync
  • Read / Write Quorum will overlap 
  • Voting does not work to read from which server
  • These systems have version numbers
  • Readers takes the ones with highest version number
  • Split database into replicas
  • Data Sharding
  • Data across protection groups
Happy Learning!!!

March 18, 2020

Distributed Systems - Session #2




I paused it a lot as I didn't really get involved much but finally managed to complete it.

Key Lessons
  • Go lang examples for threading, locking, RPC, Typesafe and memory safe, Garbage Collected
  • Threads - Tools to manage concurrency in programs
  • Stacks are within address space of the program
  • I/O Concurrency - Overlapping of progress of different activities wait ing / executing
  • Parallelism - Parallelize CPU / IO cycles / routines
  • Process is a single program / single address space. Inside process there are multiple threads
  • Process -> memory area -> routines sit inside the process
  • Process implemented by the operating system
  • Thread challenges - Sharing data
  • Mutex / Locks for shared data
  • Data Access - Managing Locks / Deadlocks / Starvation / Blocking
  • Channels (Go Lang) - Send data between threads
  • WaitGroup, Sync.Cond
  • Webcrawlers design for parallel processing using threads
  • Handling concurrency / multiple parallel threads / optimum network capacity utilization
  • Remember doing SSIS ETL parallel tasks for Data pull
A multi-threaded Web crawler implemented in Python
Crawler
Multi-Threaded Crawler in Python

Happy Learning!!!

Staying updated in Data science - My 5 Lessons

  1. Reddit, tweets, LinkedIn follows news, analytics blogs, links, Lex Fridman interviews, Stanford / MIT / Cornell updated courses
  2. Look at Kaggle kernels, understand feature variables, newer features build. Learn domain-specific findings
  3. Read research papers and try to look for techniques in video/text/ audio projects which you can reapply
  4. Look at Github examples and code them in your free time. This help to know coding practices/ best practices
  5. A lot of industry-specific products we can find by digging deep on AI technology and product landscape. Top 100 AI companies, AI product blogs, etc..
Teach, blog in different mediums. This helps to learn, gather different perspectives.  If you have observed technology and know the underlying pattern/architecture you can better connect the product, purpose, and applications of the tool.
During ML interviews I did find most interviewing folks 6 to 7 years younger than me. I came from DB BI to the AI world. It's a good feel to continue code, coach, teach a younger set of folks.

Good Read (Link)
Reading Research papers

Happy Learning!!!