January 01, 2025
April 30, 2024
Insights, Innovations, and Lessons: Exploring Computer Vision and Generative AI Hybrid solutions
Hoping to share more insights on below Vision + GenAI Journey
1. Captivating Success: Harnessing Vision and Generative AI to Mitigate Food Waste
Unveiling a remarkable deployment where vision technologies, coupled with Generative AI, are being leveraged in a significant initiative to reduce food waste. A partnership between a renowned condiment brand and a leading technology provider exemplifies the powerful application of AI in environmental sustainability.
2. Dual-edged Experiences: Enhancing Product Details and Vision Technology Setbacks
This segment will delve into the mixed outcomes from integrating Generative AI in enriching product detail pages, and the limitations encountered with computer vision technologies. We’ll share an analysis of the decision-making processes in either developing custom vision models alongside Generative AI or opting for off-the-shelf solutions, outlining the key challenges and learnings from both paths.
3. Learning from Setbacks: Challenges in Vision for Skin Care Innovations
Not all ventures yield success, and in the explorative landscape of AI, the application of vision technology for skin care solutions has faced its own set of challenges. This case will reveal the hurdles faced during implementation and the pivotal lessons learned, emphasizing the importance of iterative testing and adaptive strategies in technology application.
Summary:
In wrapping up, the session will highlight the critical takeaways from the successes and setbacks observed in integrating computer vision and Generative AI across different industries. Attendees will gain a nuanced understanding of the practical applications, scalability issues, and strategic decisions crucial for leveraging these cutting-edge technologies effectively.
Keep Exploring!!!
April 08, 2024
Anyscale Endpoints discussion
Step #1 - Anyscale signup
Step #2 - Notebook for Deploying Diffusion models
Step #3 - Deploying Service Command
Step #4 - Service Deployment
Code Example -
Keep Exploring!!!
March 20, 2024
AI - Applied use case - Vision in Action
Spot the right use case, solve with the balance of data / strategy to meet the market on time
More read - Link
Keep Exploring!!!
February 28, 2024
Video Summarization
Learning to Summarize Videos by Contrasting Clips
- Feature Extractor
- Score Predictor
- Summary Extractor
- Highlight detection as a special case of the summarization task
Video Summarization: Towards Entity-Aware Captions - Summarizing video content into a natural language description
Video Summarization Using Deep Neural Networks: A Survey
Option #1
- Feature Extractor
- Score Predictor
- Summary Extractor
- Highlight detection as a special case of the summarization task
Option #2
- Frame 1 - Feature Vector
- Frame 2 - Feature Vector2
- Frame 3 - Feature Vector 3
- Feature vector score comparison to pick / unpick
- Object score comparison to pick / unpick
Other Techniques
- Hashing based
- Clustering based
- Feature based
Keep Exploring!!!
February 07, 2024
Vision Product Catalog Startups
- Background removal
- Super resolution
- Image Restoration Deformation Fixes
Startups in Focus
Keep Exploring!!!
October 15, 2023
CNN Learning One pagers
Product and Example
- https://tangoeye.ai/
- Retail solutions built on
- Age Detection Models
- Gender Detection Models
- Face Detection
- Re-identification
- Step 1: Take a batch of training data and perform forward propagation to compute the loss.
- Step 2: Backpropagate the loss to get the gradient of the loss with respect to each weight.
- Step 3: Use the gradients to update the weights of the network.
- Chain rule derivate
- The procedure repeatedly adjusts the weights of the connections in the network so as to minimize a measure of the difference between actual output and desired output
- Ability to create new distinguishing features
- The aim is to find the set of weights that ensure that for each input vector the output vector produced by the network is same as the desired output vector
- The drawback in learning procedure is that the error surface may contain local minima so that gradient descent is not guaranteed to find a global minimum
- Introduce non-linearity into a model
- We need non-linearity, to capture more complex features and model more complex variations that simple linear models can not capture.
- neural networks use non-linear activation functions, which can help the network learn complex data, compute and learn
- Signmoid, Tanh, Relu
- The first rule of thumb is that you should not try to design your own architecture from scratch
- If you are working on generic problem, it never hurts to start with ResNet-50. If you are building a mobile-based visual application where there is limited computation resources, try MobileNets
Keep Learning!!!
August 20, 2023
Plant Identification Using Convolution Neural Network and Vision Transformer-Based Models
Recently, my team published a vision paper, providing valuable insights and lessons which will benefit our future work. Here I highlight those key experiences and challenges:
Plant Identification Using Convolution Neural Network and Vision Transformer-Based Models
- First, we grappled with open-ended questions in our problem statement, requiring us to think critically and flexibly.
- Second, we used past experiences, research approaches, and current vision models to craft our unique approach for this paper.
- Third, was the phase of experimenting which we had to analyze, timebox, and finalize.
- We also faced data challenges, drawing inspiration from similar research papers to overcome this hurdle.
- An important achievement for us was reaching state-of-art accuracy in our findings.
- We considered the scalability of our approach, contemplating how it can be implemented as we include multiple categories/classes.
- Focus was directed toward developing a repeatable architecture and effectively capturing feedback for continuous improvement.
- A significant portion of our time was dedicated to extensive documentation, conducting numerous experiments, and evaluating metrics.
- We navigated through the publication process, ensuring our work reached the right platforms.
- Lastly, we sought collaboration with like-minded clients, with whom we could work on making our learning reusable.
This experience has been thoroughly enriching for our team and we remain excited about our journey ahead
Keep Learning!!!
June 22, 2023
June 20, 2023
May 11, 2023
Vision Project Ideas
Background Remover lets you Remove Background from images and video
PyMatting: A Python Library for Alpha Matting
DeepFloyd IF, a powerful text-to-image model
Reverse image search helps you search for similar or related images given an input image.
Keep Exploring!!!
May 04, 2023
DINOv2
- Automatic pipeline to build a dedicated, diverse, and curated image dataset instead of uncurated data
- Focus on text-guided pretraining
- Textual supervision to guide the training of the features
- PCA between the patches of the images from the same column
- Features are learned from images alone
- Self-supervised learning has the potential to learn all-purposed visual features if pretrained on a large quantity of curated data
- Automatic pipeline to filter and rebalance datasets from an extensive collection of uncurated images
- Data similarities are used instead of external metadata and do not require manual annotation
Other Approaches
- Extracting a signal from the image to be predicted from the rest of the image
- Discriminative signals between images or groups of images to learn features.
- Copy detection pipeline of Pizzi et al. (2022) to the uncurated data and remove near-duplicate images
- Compute an image embedding using a self-supervised ViT-H/16 network pretrained on ImageNet-22k, and use cosine-similarity as a distance measure between images.
- k-means clustering of the uncurated data.
- Query dataset for retrieval, if it is large enough we retrieve N (typically 4) nearest neighbors for each query image.
Summary
- DINOv2, a new series of image encoders pretrained on large curated data with no supervision
- Visual features are compatible with classifiers as simple as linear layers - meaning the underlying information is readily available
Keep Exploring!!!
March 11, 2023
February 23, 2023
Startup Analysis - hyperverge - KYV - Vision + OCR + NLP
Many times taking an idea, ideating it, and solving it end to end is key. KYC with Vision / Image / Data and NLP are very impressive.
Product Features
- Real-time analysis of images and videos obtained from sources such as consumer photos, satellite images, surveillance cameras, industrial images, and documents.
- NLP solutions for automating and disrupting the Legal Document Analysis industry
Deep Learning Skills (Vision / NLP) - Our Perspectives
- DL - Models built for Face detection, OCR, and Buildings / Signs Detection. Face, Object, Text, and Activity recognition.
- NLP - One shot, Few shot, Self-Supervised approach, multi-task learning, and contrastive learning strategies
- Plus a lot of custom embeddings/graph databases/custom models
- Liveliness Check
- Background Detection
- Landmark validations
- Key facial landmarks based on submitted docs
- Similarity scores
- Social media similar image scores
- Signature Font Size, Length, Height, features of it
- Signature Font Style
- Keypoints match, landmarks match, shape, texture
- Landmarks
- Landmark distances for iris, nose, cheek, chin
- Mediapipe
- Custom Segment and measure similarity
- Classify face shapes/hairstyles
Keep Exploring!!!
February 02, 2023
Computer Vision - Startup Analysis - GAN
Computer Vision - Startup Analysis
Image Recognition - facial recognition, digital image search, image categorization, optical character recognition (OCR), defect detection, media, and security surveillance
Clarifai
- Clarifai has developed a platform for building AI-powered software solutions via API, mobile SDK Different industries,including retail, manufacturing, media, and transportation, among others
- Low-to-no-code user interface
Products
- Scribe Label - a platform for data labeling
- Spacetime Search - an AI-based visual search engine that looks for similarities
- Enlight Train - a custom model training suite packed with pre-trained AI models
- Flare - an edge AI platform that offers advanced predictive capabilities and accelerates intelligent video applications from streaming, decoding, batching
Medical Imaging
Medical imaging include x-rays, ultrasound imaging, mammography, computed tomography (CT) scans, MRI, and nuclear medicine, among others.
- Potential applications of medical imaging include confirming, assessing, treating, and following the progress of diseases and certain medical conditions
- Deep learning to assist radiologists in making data-driven decisions in breast cancer screening, treatment, and diagnosis
- AI-based breast cancer detection and treatment solutions that evaluate mammograms, leading to improved clinical results, lower healthcare costs, and better-served patients
Products
- cmTriage® is a workflow management and organization tool that sorts a radiologist’s mammogram workflow
- cmAngio® targets the detection of heart disease, leveraging AI to help doctors assess a patient’s age-adjusted risk of heart disease
- cmDensity™ cmDensity™ is a radiology tool that assists in classifying breast density
Video Analytics
CVEDIA develops and deploys AIbased video analytics for both the commercial and defense sectors.
- Deploy AI solutions for object recognition, enhanced safety, optimized efficiency, and exploring new product opportunities.
- CVEDIA-RT The company’s software offering comprises an AI-powered modular,cross-platform inference engine,providing both high and low-level interfaces.
Precision Agriculture
Precision agriculture is a management strategy incorporating high-tech tools such as sensors, GPS, robots, mapping tools, and data analytics for maximized economic return and minimized environmental impact
Analyze the data collected by hardware-agnostic sensors
Data-driven approach to crop management, empowering clients with all-encompassing analytics of their food production processes
- Yield prediction: Prospera combines machine learning and computer vision algorithms to analyze leaf-level images that help track the growth cycle and detect areas of replanting
- Field irrigation issues detection: The company uses multispectral imaging to analyze images from satellites, drones, and planes to identify areas of overwatering, underwatering, or machinery malfunctions
- Pests and disease detection: Prospera’s technology analyzes thousands of data sets from cameras installed on drones or other machinery to detect the presence and extent of certain pests or diseases.
- Crop nutritional deficiency prevention: Using Prospera’s technology, growers can get an undisrupted in-field overview and in-depth insight into problems that cannot be identified with the naked eye but are critical to crops’ nutritional value
Smart Data Capture
Smart data capture (often referred to as automated or intelligent data capture) is the process of extracting and accessing real-time data from barcodes, text, IDs, and objects by leveraging AI-based software
Scandit is a software company that develops smart device scanning solutions powered by computer vision
- The software can process up to 480 scans per minute
- It is suitable for enterprise-grade applications, featuring market leading scan accuracy and scalability to support large implementations.
- ShelfView is a cloud-based retail shelf management solution delivering data capture and analytics capabilities for greater shelf visibility and intelligent and efficient store management operations
December 02, 2022
Beauty - Paper - Research Reads / Inspirations
Paper #1 - Modifying Face Image for Ageing Marks using Specialized Filter
- shrinking image - cv2.INTER_AREA
- stretching image - cv2.INTER_CUBIC
Key Notes
- Landmarks points [0-24] represent outer face region.
- Points [25-32] represent left eyebrow region.
- Points [33-41] represent left eye region.
- Points [45-51] represent right eyebrow region.
- Points [52-59] represent right eye region.
- Points [ 68-78] represent lip regions and remaining points from [79-99] covers nose position.
- Face mask is generated using convex hull Technique [1] using the outer face landmark points.
Transformations
- We use horizontal and vertical sobel filter for detecting the wrinkles in the specific region.
- The average edge strength in each region is defined as the quantification of wrinkles feature.
- We apply a threshold condition to the edge intensity for getting the correct wrinkles from the image.
- Main concentration will be on the forehead and eye corners
Paper # - BEHOLDER-GAN: GENERATION AND BEAUTIFICATION OF FACIAL IMAGES WITH CONDITIONING ON THEIR BEAUTY LEVEL
Key Notes
- Progressive Growing of GANs (PGGAN) [13] suggested coping with the challenge of generating high-resolution images by learning first through generation of low-resolution images and progressively growing to higher resolutions
- Another important aspect of GANs is their ability to generate images with conditioning on some attribute
- CycleGAN, StarGAN
Code - Link
- This is more useful for cheek/chin expansion
- To ensure that generated image x indeed corresponds to the correct beauty level Discriminator D predict the beauty level and not just the usual real vs. fake probability
Paper - Facial makeup transfer with GAN for different aging faces
Key Notes
- Firstly removing the eyebrows and eyelashes of the input image to prepare for the eye makeup transfer
- Since the information is transferred from pixel to pixel, it needs to be fully aligned before transfering, and then layer is decomposed by the Edge Preserving Smooth Filter
Paper - Facial Makeup Transfer Combining Illumination Transfer
- Facial makeup, eye shadow and lip makeup are processed by different loss functions, and the three are integrated
- OpenCV Bilateral Filtering Algorithm [12] to achieve facial smoothing
BasicSR (Basic Super Restoration) is an open-source image and video restoration toolbox based on PyTorch
- Real-ESRGAN: A practical algorithm for general image restoration
- GFPGAN: A practical algorithm for real-world face restoration
- facexlib: A collection that provides useful face-relation functions.
- HandyView: A PyQt5-based image viewer that is handy for view and comparison.
- HandyFigure: Open source of paper figures
Paper # - Data Article template
Paper - Towards Real-World Blind Face Restoration with Generative Facial Prior
- GFPGAN consists of a degradation removal module and a pretrained face GAN as facial prior
- Facial component loss with local discriminators to further enhance perceptual facial details
- Image Restoration typically includes super-resolution, denoising, deblurring and compression removal
- Channel Split Operation is usually explored to design compact models and improve model representation ability
- latent features Flatent to map the input image to the closest latent code in StyleGAN2
- multi-resolution spatial features Fspatial for modulating the StyleGAN2 features.
Paper - SCUT-FBP: A Benchmark Dataset for Facial Beauty Perception
- We extracted 84 points as sample points containing facial contour information and shape information of the eyebrow, eyes, mouth, and so on
- The machinelearning methods we used include SVM regression (SVR), linear regression, pace regression, and Gaussian regression.
Dataset - Link
Paper - Improving Makeup Face Verification by Exploring Part-Based Representations
- The preprocessing step starts by applying a face detector followed by a 2D facial landmarks estimator, both available in DLib
- These landmarks are used to align, crop and resize the face thirds and facial parts
- left periocular, which includes the eye and eyebrow, right periocular, nose and mouth
- Makeup Face Dataset (EMFD)
- Youtube Makeup (YMU) Dataset [13]
FA-GANs: Facial Attractiveness Enhancement with Generative Adversarial Networks on Frontal Faces
- we prefer to enhance facial attractiveness via adjusting the relative distances among important facial components, such as eyes, nose, lip, and chin.
BeautyGAN - BeautyGAN: Instance-level Facial Makeup Transfer with Deep Generative Adversarial Network
Keep Exploring!!!
October 23, 2022
Metaverse, Computer Vision and Recent News
Previous article post
- Emotions tracking
- Facial Expressions
- Tiny gestures/remarks unique to the personality
- Face tracking
- More realistic presence for touch/feel senses
- Your AR / VR device is going to be enhanced to track these details
September 26, 2022
Vision Solutions - Interesting Ideas / Products / Observations / Startups
Recycle Segregation
The AMP robotic recycling system powered by an AI algorithm recognized and sorted 50 billion recycling objects, recovering all recyclables from a mixed-material stream at an accuracy of around 100 percent based entirely on image analysis.#computervision #datalabeling pic.twitter.com/pAKExe5j0K
— Keymakr (@keymakr_com) September 21, 2022
Repairs Monitoring with Vision
How #AI & #ComputerVision can be used to supervise the quality of human tasks#FutureOfWork #ArtificialIntelligence #IoT@JoannMoretti @Hana_ElSayyed @JolaBurnett @Shi4Tech @AkwyZ @CurieuxExplorer @enilev @anand_narang @labordeolivier @mvollmer1 @Fabriziobustama @PawlowskiMario pic.twitter.com/sYXId2Bb1U
— Franco Ronconi 🇮🇹 (@FrRonconi) September 14, 2022
- Number of Bolts repaired
- Duration of Repair
#ComputerVision is redefining surveillance
— Ronald van Loon (@Ronald_vanLoon) September 9, 2022
by @Seeker#AI #Drones #SmartCities #Privacy #Safety
cc: @frronconi pic.twitter.com/IuaG4hwUK4
Arms and Guns Detection
Is this a good approach?
— Charles Carter (@cctech100) September 20, 2022
US startup ZeroEyes has developed an AI-based computer vision system that detects and mitigates active shooters at schools.
Deployed at multiple sites.
More info on Superinnovators: https://t.co/TGwD19BMz6#computervision #schoolshootings #innovation pic.twitter.com/fJYAhvCuxc
Animals counting
Using AI to count Sheep 🐑🐑
— cimplify.ai (@cimplifyai) September 23, 2022
Now livestock farmers and producers can count livestock with the help of #computervision and #AI.
Source: Plainsight#technology #innovation #artificialintelligence pic.twitter.com/cggEIqexoE
I love the region of interest
- Segmentation
- Contours
- Counting
- Line of Separation
Deep Learning Market is Anticipated to Reach US$ 31.3 Billion by 2027 Registering a CAGR of 25.8%. #artificialintelligence #deeplearning #machinelearning #insight #data #tech #technology #innovation #computerscience #computervision #datascience #engineering #developers pic.twitter.com/HKfE20X1pk
— Welcome.AI (@welcomeai) August 31, 2022
Vision Startups
Cool Computer Vision Startups in 2022 https://t.co/H8GiWEEI6F via @MarkTechPost #ArtificialIntelligence #computervision #100DaysOfCode pic.twitter.com/H4YaQ2pxou
— MARKTECHPOST.COM (@Marktechpost) September 14, 2022
Keep Exploring!!!!
August 15, 2022
Summary Fashion Attributes
- Part I : Object Detection
- Part II : Attribute Tagging
- Part 3 : Recommendation based on Frequency
Fashion Meets Computer Vision: A Survey
Paper #1 - Progressive Fashion Attribute Extraction
- Attributes (neck design detailing, sleeves detailing, etc)
Paper #2 - Attr2Style: A Transfer Learning Approach for Inferring Fashion Styles via Apparel Attributes
- Low-level attributes of an apparel (for example, neck type, dress length, collar type, print etc)
- Transfer learning based approach to address the issue of style-based image captioning for a target dataset
Paper #3 - The iMaterialist Fashion Attribute Dataset
Paper #4 - A Deep-Learning-Based Fashion Attributes Detection Model
Occasion based Recommendation system in E-commerce like Amazon, Etsy
Visual Attributes for Fashion Analytics
Fine-Grained Fashion Similarity Prediction by Attribute-Specific Embedding Learning
Keep Exploring!!!




