Generative AI Leader · Free Practice Question Medium

Question 20

An online furniture retailer wants to allow customers to upload a photo of a sofa they like and ask the AI agent, "Find me a similar sofa to this, but in blue." This requires the agent to understand the content of the image and the text of the query simultaneously.

This capability is an example of what?

  • A

    Unimodal data retrieval

  • B

    Cloud Vision OCR

  • C

    Text-to-image generation

  • D

    Multimodal search

Reveal correct answer

Correct answer: D

Explanation

The query combines two different types of data (image and text) to perform a single action (search), which is the definition of multimodal search.

INCORRECT: Text-to-image generation
This involves creating a new image from a text prompt. The user is searching for existing products, not creating a new image.

CORRECT: Multimodal search
Multimodal search allows users to query using a combination of different data types (modalities), such as text, images, or audio. The agent's ability to process both the image of the sofa and the text "in blue" to find relevant products is a prime example of multimodal search.

INCORRECT: Unimodal data retrieval
This would involve searching with only one type of data, for example, searching with only text or only an image, but not combining them in a single query.

INCORRECT: Cloud Vision OCR

OCR (Optical Character Recognition) extracts text from an image. It would not be used to understand the visual characteristics (shape, style) of the sofa in the photo.

Discussion

Think the marked answer is wrong, or have a better explanation? Share it below — comments appear after review.

You must be logged in to post a comment.

Preparing For

Your Certification?

255+ certifications
Detailed explanations
Free PDF samples

Has All The Questions You Need