"No one is harder on a talented person than the person themselves" - Linda Wilkinson ; "Trust your guts and don't follow the herd" ; "Validate direction not destination" ;

March 04, 2019

Big Data Lessons #Lessons learned form Kafka in production (Tim Berglund, Confluent)

Key Lessons
Events
  • Data has to go into the database
  • Kafka - All of your data is events stream
  • Kafka is having an opinion of the world

Sensors

  • Sensor data are events
  • Car companies with Internet Connected Devices
  • Log Entries are an operational thing
  • Logs are events

Databases
  • Databases can also be events
  • Table - Collection of Key-value pairs
  • Modifications as messages
  • Updates can be stream of messages

Uses of Stream
  • Data Pipeline
  • React / Process / Transform


Event Centric Thinking
  • Web App -> Streaming Platform -> Hadoop
  • "Product View Request"
  • Forward Compatible (Receive Requests from Multiple Interfaces
  • New services can listen and easy to extend the system to evolve the system going forward

Kafka Overview
  • Producers - Kafla CLusters - Consumers
  • Data Model - Log
  • Write comes at the end
  • Log file
  • Multiple consumers can read from the log
  • Reader - Consumer
  • Writer - Producer
  • Kafka Topic = Partitioned Log
  • Kafka is Distributed Message Queue
  • Each Topic is a partitioned Log
  • Partitioned among multiple computers (Brokers)
  • The producer decides partition to write to
  • Kafka, we have ordering within the partition
  • Ordering within the partition but not available globally
  • There is no global ordering
  • Table and stream are isomorphic
  • Group of Consumers
  • Consumer groups handy way to divide among multiple consumers



Scalability of File System
  • Write Part
  • Read Part
  • Indexes, Merge, Log Tree
Kafka
  • Hundreds of MB / Sec throughput
  • Commodity hardware
  • O(1) writes
Distributed by Design
  • Replication
  • Fault Tolerance
  • Partitioning
  • Elastic Scaling
Issue #1 - Strange Happenings with Partitioning
  • Partitioning
  • One Lead Partition
  • Multiple Followers
  • One broker acts as a controller
  • Partitioning is to scale a Topic
  • Leader and Follower Partitions
  • Four Partitions and Three Replicas
  • ISR - In Sync Replica - Caught up with Leader
  • For a write to be committed it has to be commited by the leader and all other ISR
  • Watch your ISR List
  • Upgrade all the brokers in Rollout Fashion (Keep them in the same version)


Issue #2 - Automated Liveliness check
  • Broker Kept Failing
  • Leader Failure
  • More Partitions more throughput
  • More partition longer to balance cluster
  • Router Configuration problem
Issues #3 - Adding a Broker Hurts
  • Custom Environment
  • Kafka Reassignment Partition Tool
  • Generate Migration Instructions

Happy Learning Best Practices!!!

March 02, 2019

March 01, 2019

Personalized Travel Plans

I found this company Pickyourtrail interesting. I wanted to analyze more on their personalization algorithm. Some thoughts on the same.
Product Concept - Personalized End to End Travel Plans
Selling Point - Personalized recommendations
Target Audience - Well to do earning professionals

The initial dataset they collected on Travel Destinations, Trending destinations, Historical data is key

If we have to take an Automated Recommendation using ML Algorithm. I would see it this way

  • Collect Social Media Data
  • Analyze Income Range 
  • Historical Data from previous travel collected from Social Network
  • Interests collected from Likes, Comments from Social media
  • Segment the customers into categories - Wildlife Travel, Spiritual, Normad Trips, Historical Interests
  • Match their months of previous travel
  • Match based on the data collected on future interests

Possible Options - Banks can tie-up and provide such offers. Instead of credit points personalized travel plans with tie-up from these vendors
Futuristic Way - Everything tied Banks - Flight Plans - Uber - Airbnd, Everything connected end to end and adjusted based on time changes and variations

Happy Learning!!!

February 21, 2019

SVD Summary






Recommendations




























Happy Learning!!!

Analysis of MIT Deep Learning Projects

I spent sometime to Analyze the MIT Deep Learning Projects. Very Inspiring. The healthcare projects are very inspiring. Broad categories and different domains. Good Read to know use cases and architecture.


Updated link











Happy Mastering DL!!!

Segmentation of Data Scientists

Data Scientists from stats world - This cluster has PhDs from the 2000s and working in Vision, Analytics since 2K period. Conversations with them were useful to handcraft features for image processing problems. They know the algos, basic math involved, intuitive details and the limitations of techniques.

Data Scientists with domain expertise - Laterals upskilled with data science skills. Data science practitioner world. Ability to bridge domain and Data Science use cases. Their Strength lies in identifying data, building the pipeline. Envisioning the end to end use flow.

Rookies - These days MOOC, Coursera, Udemy, Online Sessions, data science has a lot of visibility and attention for Entry level career choice. A lot of entry-level folks getting deeper into building models, getting good at model building, feature engineering

Kaggle Experts - The goto guys on feature engineering, parameter tuning, experimenting models, applying ensemble techniques, build models from anonymized data with the best accuracy

My journey has been through Databases, BI, Analytics. I use database primarily to data analysis, the perspective of BI helps to understand the Data from the business context, domain knowledge helps to quickly extract key data and quickly build models. All this experience helps to find use cases, building features for data models, build the data model, and sell it to business. I am still getting better in *selling part*. I keep learning with my interactions from all the segments of Data Scientists

Updated - 2022 - Feb 21


Ref - Link


Happy Mastering DL!!!


February 19, 2019