"No one is harder on a talented person than the person themselves" - Linda Wilkinson ; "Trust your guts and don't follow the herd" ; "Validate direction not destination" ;
Showing posts with label Vision. Show all posts
Showing posts with label Vision. Show all posts

March 24, 2025

AI Calories Scam

  • Near Pose - X Calories
  • Far Pose - X/2 Calories
  • Side Pose - 3X Calories



Ref - Link

Keep Thinking!!!

March 06, 2025

November 21, 2024

Old Ad vs GenAI Ad

Coco Cola Old Ad



Coco Cola GenAI Ad




Keep Thinking!!!

November 05, 2024

Vision Use case

How to Implement the Use Case Correctly

  • Field of View
  • Stable Infrastructure
  • Minimal Occlusion
  • No Manual Calibration
  • With a good setup, half of the complexity and noise can be eliminated.

Keep Exploring!!!

March 28, 2024

AI in Vision Marketing

My post last year


AI generated Ad

Keep Exploring!!!


February 13, 2024

Good Read - Three trends of 2024 - Multimodal race

Good Read - Link

  • Small models - high-quality training data. Possible use cases - LLM on edge with high accuracy
  • Multimodal AI - Merge text + image + video + Audio. Making all types of content useful for creating/querying new assets. Creative Breakthru across image/video / ads/games etc
  • AI in healthcare/agriculture

For #2 - Multimodal race has a lot of players

  • Winning products across all modalities 
  • Platforms that enable creating + publishing content workflows
  • Automatic content creation, customization, and repurposing across formats and platforms

microsoft designer is stunning. 

#2024 #predictions

Keep Exploring!!!


February 04, 2024

Computer Vision License Validation

Business problem: Id verification system(valid or invalid) say driving license as id. How do we go about solving this business problem using Deep learning

Input - License Id Images

Approach

  • Feature Definition
  • Defining Elements
  • Historical data
  • Labeling / Annotation

Vision

  • Problem #1 - Extract Face images
  • Problem #2 - OCR, License Id, Dates, LicenseNumber, Authority
  • Problem #3 - Detection for Signature Extracting
  • Data Validation - Blurriness - Image Sharpening / Laplacian / Sobel / Canny edge to sharpen images. Non-readable - Far / Validation - Near View

Backend Validation

  • API call
  • Face Match
  • Similarity Score
  • Output - Valid License

Keep Exploring!!!

January 30, 2024

Startup Analysis - Jan20 - Concepting Tools in Vision - dreamlook.ai

Concepting, Vision based ideation is picking up. Dreambooth custom fine-tuning is easy to use, intuitive, and user-friendly. The UX and execution are seamless. 

Below are steps in a basic working example

1. Start with Custom Model Training


2. Upload Images
3. Submit a Job

4. Monitor Job Progress
5. Job Completion
6. Generate Images based on Custom Models


Very user friendly tool.

Keep Exploring!!!

January 18, 2024

AI Tools + Vision Use Cases + GenAI

Vision Tools + GenAI

  • Stable Diffusion, ComfyUI and Automatic1111.
  • Dreambooth and LoRA
  • Midjourney, Dalle, Runway, and PikaLabs
  • Supportive AI tools for segmentation, data labelling and inspection
  • NeRFs and Gaussian Splatting
  • DALL-E, Runway and Wonder Studio

Use Cases

  • Commercial Production
  • Graphic Design
  • Social Media
  • Content Marketing
  • Branding
  • Product Mockups
  • Spec Ads

Domain-Specific Use Cases

  • Drafting concept art, architectural concepts, and interior design plans on a budget
  • Generating free portraits of yourself, friends, family members, and pets
  • Completing hand-drawn projects that you no longer have free time for
  • Designing stunning cover art for podcasts, albums, and books
  • Printing AI-generated posters that fit your aesthetic
  • Crafting custom gifts for birthdays and holidays
  • Generating wallpapers and backgrounds for your desktop or phone
  • Visualizing random ideas to get your creativity flowing
  • Mixing up your social media posts with a new style
  • Writing cards and invitations for personal and commercial use
  • Creating eye-catching clipart-style characters for emails, posts, and presentations
  • Developing logos and icons for websites, apps, and marketing
  • Experimenting with fashion design projects
  • Competing in art challenges to embrace the AI community
  • Growing your business with AI art prints
Keep Exploring!!!

January 08, 2024

GenAI based AI interior tools

Very interesting prototype and design choices from post

Key Learning's

  • 360 view of the room
  • GenAI based image inpainting
  • By using a 360° photo of the space we can move around in different positions and sizes to have a preview of each section, and we can fully expand the space to have a 360° immersive experience as shown below.
  • Generative AI is inpainting and outpainting
  • Regenerate variants to view more options
Keep Exploring!!!

January 04, 2024

January 03, 2024

Vision Learning Startup - curiousrefuge

curiousrefuge

Made an Adidas AI Spec Commercial during Coffee Break
byu/Theblasian35 inmidjourney

The World’s First Home for AI Storytellers 

Ai-filmmaking

  • Ideation + Scriptwriting
  • Art Direction + Curation
  • Prompt Mastering + Directing
  • Pitching + Storyboarding
  • Editing + Pacing + Character
  • Cinematography + VFX

AI Filmmaking Tools

The Best Text-to-Image AI Tool

  • Midjourney
  • Adobe Firefly
  • Dall-E 3
  • Leonardo
  • Stable Diffusion

The Best Image-to-Video AI Tool

  • Runway Gen 2
  • Pika Labs

Text-to-Video AI Tool

  • Runway Gen 2
  • Moonvalley

Language Processing: Script Development, Outlining, Research, Distribution, & More. ChatGPT4

The Best AI Music for Filmmakers Tool

  • Soundful
  • Stable Audio 
  • Google MusicML 

AI Text-to-Voice Tool for Filmmakers - Elevenlabs

Best AI Tool for Voice Cloning - Elevenlabs

Best AI Tool for Generative Inpainting - Midjourney Inpainting

Keep Exploring!!!

November 30, 2023

The Future of Product Innovation is Here! 🌟 #NextGenIdeas #CollaborativeMagic

  • 🎯 Go niche with a Domain-Specific GPT. Why settle for generic when you can bring your own data and craft a model that knows your field inside out? #CustomizedAI #FocusedExcellence
  • 💡 It's a Multimodal World! Knowledge isn't just text; it's Images + Text + Data. Embrace the power of combined data forms to receive enriched, multimodal insights that tell the complete story. #HolisticAI #MultimodalKnowledge
  • 🚀 Supercharge your innovation engine with Ideas an ensemble of Models. Watch your prototyping speed take off as diverse AI models converge to refine your visions faster than ever! #SpeedyPrototyping #AIEnsemble
  • 🛡️ Wave goodbye to one-size-fits-all solutions and hello to Personal Data + Private LLM. Your privacy remains intact while you enjoy tailor-made advice crafted just for you. #PersonalizedAdvice #PrivacyMatters
  • 🛠️ Be part of the vanguard shaping AI Infrastructures. It's not just about building models; it's about building them right—robust, reliable, and fair. For all infrastructure providers, evaluators, and advocates for responsible AI, your insights are invaluable. #AIEthics #ResponsibleAI
  • Let's merge creativity with cutting-edge tech to craft products that not only resonate but revolutionize. Your privacy, our innovation—let's create the future, together.
Make the leap into a new era of product development and #TransformativeCollaboration today! 

Keep Exploring!!!

November 12, 2023

Vertex AI Vision

Some key steps to experiment in coming weeks. This low-code vision platform has been in my to-do list. Bookmarking some references

Stream registration

Open the Streams tab of the Vertex AI Vision dashboard.

  1. Go to the Streams tab
  2. Click addRegister.
  3. Enter the stream name and select a region. You can click Add Row to register multiple streams at the same time.
  4. Click the Register button to create one or more streams.

# This command streams a video file to a stream. Streaming ends when the video ends.
vaictl -p PROJECT_ID \
         -l LOCATION_ID \
         -c application-cluster-0 \
         --service-endpoint visionai.googleapis.com \
send video-file to streams STREAM_ID --file-path LOCAL_FILE.EXT

# This command streams a video file to a stream. Video is looped into the stream until you stop the command.
vaictl -p PROJECT_ID \
         -l LOCATION_ID \
         -c application-cluster-0 \
         --service-endpoint visionai.googleapis.com \
send video-file to streams STREAM_ID --file-path LOCAL_FILE.EXT --loop

export SOURCE=gs://cloud-samples-data/vertex-ai-vision/street_vehicles_people.mp4
gsutil cp $SOURCE .

export PROJECT_ID=<Your Google Cloud project ID>
export LOCATION_ID=us-central1
export LOCAL_FILE=street_vehicles_people.mp4

nohup vaictl -p $PROJECT_ID \
    -l $LOCATION_ID \
    -c application-cluster-0 \
    --service-endpoint visionai.googleapis.com \
send video-file to streams 'traffic-stream' --file-path $LOCAL_FILE --loop &


Keep Exploring!!!


October 29, 2023

Google Vision - Experiments - Vertex Vision

  • Cloud Video Intelligence API - Detects objects, explicit content, and scene changes in videos. It also specifies the region
  • Cloud Vision API - Image Content Analysis

Track objects in a streaming video

Track objects

Shot change detection tutorial

  • SHOT_CHANGE_DETECTION request 
  • List of all shots that occur within the video
  • For each shot, provide the start and end time of the shot

Track objects in a local video file

All Video Intelligence code samples

AI-powered video archive for searching family videos

Video intelligence takes to the streets



Hello video data: Train an AutoML video classification model

Live Streaming on Google Cloud with Media CDN and Live Streaming API

Experiments

  • On lower resolution
  • Out-of-box detections
  • Select frames by Shot detection and evaluate
  • Offline mode evaluation (On videos from the bucket)
  • Online mode evaluation (Live Streaming API)
  • To find unknowns/limitations from sample videos 
  • Exploratory analysis/video/image

Architecture options

  • Offline Video evaluation with video in GCP bucket
  • Offline Video evaluation + custom models
  • Offline Video + Shot Detection + Out of box object detection
  • Live Streaming API Evaluation




Register streams - Streams connect your physical devices (like IP cameras)
Create Apps


Keep Exploring!!!

October 23, 2023

Text to Vision - Image - Survey - Techniques - Lessons

Text to Vision - Image - Survey - Techniques - Lessons

  • multimodal-to-text generation models (e.g. Flamingo)
  • image-text matching models (e.g. CLIP)
  • text-to-image generation models (e.g. Stable Diffusion).

A Survey of Diffusion Based Image Generation Models: Issues and Their Solutions

  • Difficulty generating images with multiple objects
  • Quality improvement of generated images

Concepts

  • To generate images with multiple objects, layout information such as bounding boxes or segmentation maps is added to the model. 
  • Cross-attention maps have been found to play a crucial role in image generation quality
  • Techniques like “SynGen” [55] and “Attend-and-Excite” [9] have been introduced to improve attention maps

  • Mathematically, this process can be modeled as a Markov process.
  • The process of adding noise step by step from X0 to XT is called “forward process” or “diffusion process”
  • Reversely, from XT , the process of iterative remove the noise until getting the clear image is called the “reverse process”

  • Denoising Diffusion Probabilistic Models (DDPM)
  • Basic components of a diffusion model
  • Noise prediction module - U-net / pure transformer structure
  • Condition encoder - conditioned on something, such as text. T5 series encoder or CLIP text encoder is used in most of the current works.
  • Super resolution module - DALL·E 2 employs two super-resolution models in its pipeline
  • Dimension reduction module - Text encoder and image encoder of CLIP are components integrated into the DALL·E 2 model
  • Diffusion models can also encounter difficulties in accurately representing positional information
  • SceneComposer [75] and SpaText [1] concentrate on leveraging segmentation maps for image synthesis
  • Subject Driven Generation
  • Concept customization or personalized generation
  • Present an image or a set of images that represent a particular concept, and then generate new images based on that specific concept




  • Advantage of Blip-diffusion lies in its ability to perform “zero-shot” generation, as well as “few-shot” generation with minimal fine-tuning

QUALITY IMPROVEMENT OF GENERATED IMAGES

  • Mixture of experts (MOE) [60] is a technique that leverages the strengths of different models, and it has been adapted for use in diffusion models to optimize their performance
  • Employ Gaussian blur on the on certain area of the prediction according to self-attention map to extract this condition

Reverse Stable Diffusion: What prompt was used to generate this image?

  • new task of predicting the text prompt given an image generated by a generative diffusion model
  • DiffusionDB is the first large-scale text-to-image prompt dataset. It contains 14 million images generated by Stable Diffusion using prompts and hyperparameters specified by real users.
  • Diffusion Explorer

Learning framework for prompt embedding estimation

Reversing the textto-image diffusion process

  • Predict a sentence embedding of the original prompt used to generate the input image
  • As underlying models, we consider three state-of-the-art architectures
  • that are agnostic to the generative mechanism of Stable Diffusion, namely ViT, CLIP and Swin Transformer 
  • U-Net model from Stable Diffusion, which operates in the latent space.

Explain in laymen terms - U-Net model from Stable Diffusion, which operates in the latent space.

The U-Net model from Stable Diffusion is a type of artificial intelligence model used for various computer vision tasks like image segmentation, where it identifies and separates different objects or features within an image.

Imagine you have a picture that appears blurry, full of noise, or unclear. The U-Net model from Stable Diffusion operates like a sophisticated visual detective, which can work back through the noise, step by step, to try and reconstruct the original picture.

To do this, it operates in what we call the 'latent space', which is loosely analogous to the mind’s eye of the AI - it's where the AI forms a sort of abstract, compressed understanding of the different elements present in the image, their shapes, and how they relate to each other. You can think of the latent space as a box where the details of the image are stored in a compact form, almost like the raw components before they've been assembled into the complete picture.

So, the U-Net model from Stable Diffusion first takes a noisy image, maps or translates it into this intermediate latent space - compressing and organizing the information in a way it can handle - before then reconstructing the original, clearer image from that. It's essentially a way of moving from a jumble of details,into a structured "blueprint" in the latent space, and then using that blueprint to rebuild a clear and accurate image. 

A key aspect of the U-Net model is its structure, which is like a U-shape (thus the name 'U-net'). The first half of the U shape takes the noisy image and condenses it down into the blueprint in the latent space (this is called encoding or downsampling). The second half then expands this blueprint back out into the clear image (known as decoding or upsampling). This U-shape structure, combined with the operation in the latent space, allows the model to effectively manage and recover the important details from the noisy input and improve the generated output's quality significantly.

So in simple terms, the U-Net model from Stable Diffusion operates like a skilled restorer, turning a distorted or noisy picture back into a clear and identifiable image by operating in its “mind’s eye” or latent space, using a special U-shaped structure to carefully manage detail extraction and restoration.


Diffusion Explainer: Visual Explanation for Text-to-image Stable Diffusion


VLP: A Survey on Vision-Language Pre-training

Image Feature Extraction

  • By using the Faster R-CNN, VLP models obtain the OD-based Region feature embedding
  • CNNs end-to-end by using the grid features 

Video Feature Extraction

  • VLP models [17, 18] extract the frame features by using the method mentioned above

Text Feature Models

  • For the textual features, following pretrained language model such as BERT [2], RoBERTa [24], AlBERT [25], and XLNet [26], VLP models [9, 27, 28] first segment the input sentence into a sequence of subwords



Ideas summary

  • Detection models for Region feature embedding
  • Grid based feature extraction with CNN
  • Super resolution module to the pipeline
  • Subject Driven Generation, Concept customization or personalized generation - Present an image or a set of images that represent a particular concept, and then generate new images based on that specific concept
  • Gaussian blur on the on certain area based on attention / relevance
  • Captioning, Category recognition
  • Category Recognition (CR) CR refers to identifying the category and sub-category of a product, such as {HOODIES, SWEATERS}, {TROUSERS, PANTS}
  • Multi-modal Sentiment Analysis (MSA) MSA is aimed to detect sentiments in videos by leveraging multi-modal signals 

Text-to-image Diffusion Models in Generative AI: A Survey

The learning goal of DM is to reserve a process of perturbing the data with noise, i.e. diffusion, for sample generation

Diffusion Probabilistic Models (DPM), Score-based Generative model(SGM)

Denoising diffusion probabilistic models (DDPMs) are defined as a parameterized Markov chain

  • Forward pass. In the forward pass, DDPM is a Markov chain where Gaussian noise is added to data in each step until the images are destroyed
  • Reverse pass. With the forward pass defined above, we can train the transition kernels with a reverse process

Conditional diffusion model: A conditional diffusion model learns from additional information (e.g., class and text) by taking them as model input.

Guided diffusion model: During the training of a guided diffusion model, the class-induced gradients (e.g. through an auxiliary classfier) are involved in the sampling process.


Awesome Video Diffusion

Keep Exploring!!!

Interesting Product - sivi.ai

sivi.ai

The concept of blending image/text and providing ad variations is very impressive :)




The next question comes up / How does it compete against other models? Text to image generator options?




Current State of Art models struggle with creating the right mix of design with image and text content.

My Understanding

  • Have variations for text
  • Have Variations for image
  • Leverage past data
  • Position according to domain/data
  • Generate variations


30% image variations
30% text variations
40% templates and positioning based on domain / data / templates

Keep Exploring!!!

October 20, 2023

Google Vertex Vision - Analytics - GenAI - Vertex Matching Engine

Vertex Vision 

Feed real-time streaming video

Pick existing models

Plug custom vision models


Architecture references and GenAI - Vision + Text + Catalog Management

Summary items

  • Finetuning with sample images - Few shot learning
  • Step 1 - Image Embedding Extractor - Vertex AI Embedding Extractor to extract embedding for image
  • Step 2 - Vertex Matching Engine to fetch top and similar images
  • Step 3 - Create a new copy for images and new Text, and Upload to the product database

Pre-requisites - Catalog of images


Ref - Accelerate product catalog management with generative AI 

Step 1 - Embedding Extract 

Step 2 - Similar Products

Step 3 - Add Descriptions

Step 4 - Prompt based enrichment

Advantage - Language translations supported

Step 5 - Catalog image creation


Ref -  Accelerating product innovation with generative AI

Step 1- Text data import - reviews, product info

Step 2 - Extract insights from uploaded info

Step 3 - QnA

Step 4 - Product Generation from concept

Summary

  • Concept one-liner (1 word)
  • Features from concept details (Few lines)
  • Prepare description with features (Product V1 Description)
  • Description to create Images (Image template creation)
  • Inspiration with details and images (Draft Product Ready)

Keep Exploring!!!