- Near Pose - X Calories
- Far Pose - X/2 Calories
- Side Pose - 3X Calories
Ref - Link
Keep Thinking!!!
Deep Learning - Machine Learning - Data(base), NLP, Video - SQL Learning's - Startups - (Learn - Code - Coach - Teach - Innovate) - Retail - Supply Chain
Coco Cola Old Ad
Back in 1995, Coca-Cola launched their famous “Holidays Are Coming” ad.
— Salma (@Salmaaboukarr) November 20, 2024
It showed a convoy of Christmas-lit trucks arriving in a snowy town, spreading joy.
It became a classic and made Coca-Cola a big part of holiday traditions. pic.twitter.com/S5KAi5SqgM
What is Generative AI?
— Salma (@Salmaaboukarr) November 20, 2024
It’s advanced tech that creates text, images, audio, and video.
Coca-Cola used it to:
• Capture the nostalgic essence of the 1995 ad
• Reimagine it with modern visuals
• Create something that couldn’t have been done before pic.twitter.com/tDvmGElFUM
How to Implement the Use Case Correctly
My post last year
AI generated Ad
Keep Exploring!!!
Some moments to cherish :)
Hellmann’s collaborates with Google on AI tool that tackles food waste
Hellmann’s Launches Innovative Campaign to Clear the Galaxy of Food Waste
SANDWICH HELLMANNS
Happy Learning!!!
Good Read - Link
For #2 - Multimodal race has a lot of players
microsoft designer is stunning.
#2024 #predictions
Keep Exploring!!!
Business problem: Id verification system(valid or invalid) say driving license as id. How do we go about solving this business problem using Deep learning
Input - License Id Images
Approach
Vision
Backend Validation
Concepting, Vision based ideation is picking up. Dreambooth custom fine-tuning is easy to use, intuitive, and user-friendly. The UX and execution are seamless.
Below are steps in a basic working example
1. Start with Custom Model Training
Keep Exploring!!!
Vision Tools + GenAI
Use Cases
Domain-Specific Use Cases
Very interesting prototype and design choices from post
Key Learning's
Made an Adidas AI Spec Commercial during Coffee Break
byu/Theblasian35 inmidjourney
The World’s First Home for AI Storytellers
The Best Text-to-Image AI Tool
The Best Image-to-Video AI Tool
Text-to-Video AI Tool
Language Processing: Script Development, Outlining, Research, Distribution, & More. ChatGPT4
The Best AI Music for Filmmakers Tool
AI Text-to-Voice Tool for Filmmakers - Elevenlabs
Best AI Tool for Voice Cloning - Elevenlabs
Best AI Tool for Generative Inpainting - Midjourney Inpainting
Keep Exploring!!!
Some key steps to experiment in coming weeks. This low-code vision platform has been in my to-do list. Bookmarking some references
Open the Streams tab of the Vertex AI Vision dashboard.
export SOURCE=gs://cloud-samples-data/vertex-ai-vision/street_vehicles_people.mp4
gsutil cp $SOURCE .
export PROJECT_ID=<Your Google Cloud project ID>
export LOCATION_ID=us-central1
export LOCAL_FILE=street_vehicles_people.mp4
nohup vaictl -p $PROJECT_ID \
-l $LOCATION_ID \
-c application-cluster-0 \
--service-endpoint visionai.googleapis.com \
send video-file to streams 'traffic-stream' --file-path $LOCAL_FILE --loop &
Track objects in a streaming video
Shot change detection tutorial
Track objects in a local video file
All Video Intelligence code samples
AI-powered video archive for searching family videos
Video intelligence takes to the streets
Hello video data: Train an AutoML video classification model
Live Streaming on Google Cloud with Media CDN and Live Streaming API
Experiments
Architecture options
Text to Vision - Image - Survey - Techniques - Lessons
A Survey of Diffusion Based Image Generation Models: Issues and Their Solutions
Concepts
QUALITY IMPROVEMENT OF GENERATED IMAGES
Reverse Stable Diffusion: What prompt was used to generate this image?
Learning framework for prompt embedding estimation
Reversing the textto-image diffusion process
Explain in laymen terms - U-Net model from Stable Diffusion, which operates in the latent space.
The U-Net model from Stable Diffusion is a type of artificial intelligence model used for various computer vision tasks like image segmentation, where it identifies and separates different objects or features within an image.
Imagine you have a picture that appears blurry, full of noise, or unclear. The U-Net model from Stable Diffusion operates like a sophisticated visual detective, which can work back through the noise, step by step, to try and reconstruct the original picture.
To do this, it operates in what we call the 'latent space', which is loosely analogous to the mind’s eye of the AI - it's where the AI forms a sort of abstract, compressed understanding of the different elements present in the image, their shapes, and how they relate to each other. You can think of the latent space as a box where the details of the image are stored in a compact form, almost like the raw components before they've been assembled into the complete picture.
So, the U-Net model from Stable Diffusion first takes a noisy image, maps or translates it into this intermediate latent space - compressing and organizing the information in a way it can handle - before then reconstructing the original, clearer image from that. It's essentially a way of moving from a jumble of details,into a structured "blueprint" in the latent space, and then using that blueprint to rebuild a clear and accurate image.
A key aspect of the U-Net model is its structure, which is like a U-shape (thus the name 'U-net'). The first half of the U shape takes the noisy image and condenses it down into the blueprint in the latent space (this is called encoding or downsampling). The second half then expands this blueprint back out into the clear image (known as decoding or upsampling). This U-shape structure, combined with the operation in the latent space, allows the model to effectively manage and recover the important details from the noisy input and improve the generated output's quality significantly.
So in simple terms, the U-Net model from Stable Diffusion operates like a skilled restorer, turning a distorted or noisy picture back into a clear and identifiable image by operating in its “mind’s eye” or latent space, using a special U-shaped structure to carefully manage detail extraction and restoration.
Diffusion Explainer: Visual Explanation for Text-to-image Stable Diffusion
VLP: A Survey on Vision-Language Pre-training
Image Feature Extraction
Video Feature Extraction
Text Feature Models
Ideas summary
Text-to-image Diffusion Models in Generative AI: A Survey
The learning goal of DM is to reserve a process of perturbing the data with noise, i.e. diffusion, for sample generation
Diffusion Probabilistic Models (DPM), Score-based Generative model(SGM)
Denoising diffusion probabilistic models (DDPMs) are defined as a parameterized Markov chain
Conditional diffusion model: A conditional diffusion model learns from additional information (e.g., class and text) by taking them as model input.
Guided diffusion model: During the training of a guided diffusion model, the class-induced gradients (e.g. through an auxiliary classfier) are involved in the sampling process.
Keep Exploring!!!
The concept of blending image/text and providing ad variations is very impressive :)
The next question comes up / How does it compete against other models? Text to image generator options?
My Understanding
Feed real-time streaming video
Pick existing models
Plug custom vision models
Architecture references and GenAI - Vision + Text + Catalog Management
Summary items
Pre-requisites - Catalog of images
Ref - Accelerate product catalog management with generative AI
Step 1 - Embedding Extract
Step 2 - Similar Products
Step 3 - Add Descriptions
Step 4 - Prompt based enrichment
Advantage - Language translations supported
Step 5 - Catalog image creation
Ref - Accelerating product innovation with generative AI
Step 1- Text data import - reviews, product info
Step 2 - Extract insights from uploaded info
Step 3 - QnA
Step 4 - Product Generation from concept
Summary
Keep Exploring!!!
For questions/feedback/career opportunities/training / consulting assignments/mentoring - please drop a note to sivaram2k10(at)gmail(dot)com
Coach / Code / Innovate