Multimodal Support
DSPy.rb can pass images and supported PDF documents through raw chat or typed signatures. Provider and input-shape limits differ, so choose the adapter before designing the signature.
Vision-Capable Models
OpenAI Models
gpt-4oandgpt-4o-mini
Anthropic Models
- Claude 4 series (Opus, Sonnet)
- Claude 3.5 series (Sonnet, Haiku)
Google Gemini Models
gemini-2.5-flashgemini-2.5-pro
PDF Document Support
PDF document inputs currently have a narrower contract than images:
DSPy::Documentsupports PDF documents only (application/pdf)- Direct document support works with Anthropic models
- RubyLLM document support works only when the underlying provider is Anthropic
Predictsupports exactly one top-level document input per call- Mixed image and document
Predictinputs are not supported in this release
Creating Documents
# From URL
document = DSPy::Document.new(
url: 'https://example.com/report.pdf'
)
# From base64 data
document = DSPy::Document.new(
base64: pdf_base64,
content_type: 'application/pdf'
)
# From byte data
File.open('report.pdf', 'rb') do |file|
document = DSPy::Document.new(
data: file.read.bytes,
content_type: 'application/pdf'
)
end
Using Documents with Raw Chat
lm = DSPy::LM.new('anthropic/claude-sonnet-4-20250514', api_key: ENV['ANTHROPIC_API_KEY'])
document = DSPy::Document.new(
url: 'https://example.com/report.pdf'
)
response = lm.raw_chat do |messages|
messages.system('You are a financial analyst.')
messages.user_with_document('Summarize the key metrics in this PDF.', document)
end
Using Documents with Predict
Predict can attach one top-level DSPy::Document input and preserve a placeholder in the rendered prompt:
class DocumentSummary < DSPy::Signature
description 'Extract a summary from a PDF document'
input do
const :document, DSPy::Document, description: 'PDF document to summarize'
const :focus, String, description: 'What to focus on'
end
output do
const :summary, String, description: 'Document summary'
end
end
predictor = DSPy::Predict.new(DocumentSummary)
result = predictor.call(
document: document,
focus: 'financial metrics'
)
puts result.summary
Using Documents Through RubyLLM
lm = DSPy::LM.new('ruby_llm/claude-sonnet-4-5', api_key: ENV['ANTHROPIC_API_KEY'])
response = lm.raw_chat do |messages|
messages.user_with_document('Extract the revenue numbers.', document)
end
Create Images
Creating Images
DSPy::Image accepts a URL, base64 data, or byte data:
# From URL (OpenAI only)
image = DSPy::Image.new(
url: 'https://example.com/image.jpg'
)
# From base64 data (both providers)
image = DSPy::Image.new(
base64: 'iVBORw0KGgoAAAANSUh...', # your base64 string
content_type: 'image/jpeg'
)
# From byte array (both providers)
File.open('image.jpg', 'rb') do |file|
image = DSPy::Image.new(
data: file.read,
content_type: 'image/jpeg'
)
end
# With detail level (OpenAI only)
image = DSPy::Image.new(
url: 'https://example.com/image.jpg',
detail: 'high' # 'low', 'high', or 'auto'
)
Supported Formats
- JPEG (
image/jpeg) - PNG (
image/png) - GIF (
image/gif) - WebP (
image/webp)
Size Limits
- Maximum size: 5MB per image
- Multiple images: Supported (counts toward token usage)
Using Images with LM
Simple Image Analysis
# Initialize with a vision-capable model
lm = DSPy::LM.new('openai/gpt-4o-mini', api_key: ENV['OPENAI_API_KEY'])
# Create an image
image = DSPy::Image.new(url: 'https://example.com/photo.jpg')
# Analyze the image
response = lm.raw_chat do |messages|
messages.user_with_image('What is in this image?', image)
end
puts response
Multiple Images
image1 = DSPy::Image.new(url: 'https://example.com/before.jpg')
image2 = DSPy::Image.new(url: 'https://example.com/after.jpg')
response = lm.raw_chat do |messages|
messages.user_with_images(
'What changed between these two images?',
[image1, image2]
)
end
With System Prompts
response = lm.raw_chat do |messages|
messages.system('You are an expert art critic.')
messages.user_with_image('Analyze this painting.', image)
end
Structured Multimodal Signatures
Structured signatures can declare typed outputs for image analysis.
Image Analysis with Structured Output
This signature extracts colors, objects, mood, and style:
class ImageAnalysis < DSPy::Signature
description "Extract objects, colors, mood, and style from images"
class DetailLevel < T::Enum
enums do
Brief = new('brief')
Standard = new('standard')
Detailed = new('detailed')
end
end
input do
const :image, DSPy::Image, description: 'Image to analyze'
const :focus, String, default: 'general', description: 'Analysis focus'
const :detail_level, DetailLevel, default: DetailLevel::Standard, description: 'Level of detail'
end
output do
const :description, String, description: 'Overall description of the image'
const :objects, T::Array[String], description: 'List of objects detected'
const :dominant_colors, T::Array[String], description: 'Main colors in the image'
const :mood, String, description: 'Overall mood or atmosphere'
const :style, String, description: 'Artistic style or characteristics'
const :lighting, String, description: 'Description of lighting conditions'
const :confidence, Float, description: 'Analysis confidence (0.0-1.0)'
end
end
# Usage
analyzer = DSPy::Predict.new(ImageAnalysis)
image = DSPy::Image.new(url: 'https://example.com/landscape.jpg')
analysis = analyzer.call(
image: image,
focus: 'colors',
detail_level: ImageAnalysis::DetailLevel::Detailed
)
puts analysis.description
puts "Colors: #{analysis.dominant_colors.join(', ')}"
puts "Mood: #{analysis.mood}"
puts "Objects: #{analysis.objects.join(', ')}"
Object Detection with Type-Safe Bounding Boxes
Use T::Struct for type-safe bounding box detection:
# Define structured types
class BoundingBox < T::Struct
const :x, Float
const :y, Float
const :width, Float
const :height, Float
end
class DetectedObject < T::Struct
const :label, String
const :bbox, BoundingBox
const :confidence, Float
end
class BoundingBoxDetection < DSPy::Signature
description "Detect and locate objects in images with normalized bounding box coordinates"
class DetailLevel < T::Enum
enums do
Basic = new('basic')
Standard = new('standard')
Detailed = new('detailed')
end
end
input do
const :query, T.any(String, NilClass), description: 'Object to detect'
const :image, DSPy::Image, description: 'Image to analyze for object detection'
const :detail_level, DetailLevel, default: DetailLevel::Standard, description: 'Detection detail level'
end
output do
const :objects, T::Array[DetectedObject], description: 'Detected objects with bounding boxes'
const :count, Integer, description: 'Total number of objects detected'
const :confidence, Float, description: 'Overall detection confidence'
end
end
# Usage with type safety
detector = DSPy::Predict.new(BoundingBoxDetection)
image = DSPy::Image.new(url: 'https://example.com/aerial-image.jpg')
detection = detector.call(
query: 'airplanes',
image: image,
detail_level: BoundingBoxDetection::DetailLevel::Standard
)
detection.objects.each do |obj|
puts "#{obj.label} at (#{obj.bbox.x}, #{obj.bbox.y})"
puts "Size: #{obj.bbox.width} x #{obj.bbox.height}"
puts "Confidence: #{(obj.confidence * 100).round(1)}%"
end
Working with Anthropic Models
When using Anthropic models, you need to provide images as base64 or raw data:
# Configure Anthropic model
lm = DSPy::LM.new('anthropic/claude-4', api_key: ENV['ANTHROPIC_API_KEY'])
# Load and encode image as base64
File.open('image.jpg', 'rb') do |file|
image_data = file.read
base64_data = Base64.strict_encode64(image_data)
image = DSPy::Image.new(
base64: base64_data,
content_type: 'image/jpeg'
)
response = lm.raw_chat do |messages|
messages.system('You are an image analysis expert.')
messages.user_with_image('Describe this image in detail.', image)
end
puts response
end
Send Images to Google Gemini Models
The Gemini adapter accepts base64 or byte image data:
# Configure Gemini model
lm = DSPy::LM.new('gemini/gemini-2.5-flash', api_key: ENV['GEMINI_API_KEY'])
# Load and encode image as base64
File.open('product_image.jpg', 'rb') do |file|
image_data = file.read
base64_data = Base64.strict_encode64(image_data)
image = DSPy::Image.new(
base64: base64_data,
content_type: 'image/jpeg'
)
response = lm.raw_chat do |messages|
messages.system('You are a product analysis expert.')
messages.user_with_image('Analyze this product image for e-commerce listing.', image)
end
puts response
end
Multiple Images with Gemini
# Analyze multiple product angles
images = ['front.jpg', 'back.jpg', 'side.jpg'].map do |filename|
File.open(filename, 'rb') do |file|
DSPy::Image.new(
base64: Base64.strict_encode64(file.read),
content_type: 'image/jpeg'
)
end
end
response = lm.raw_chat do |messages|
messages.user_with_images(
'Compare these product images and identify any defects or quality issues.',
images
)
end
puts response
Platform Differences
OpenAI
- URL Support: Direct URL references supported
- Detail Levels: Can specify
low,high, orautodetail - Token Costs: Images consume tokens based on size and detail
Anthropic
- Base64 Only: Images must be base64-encoded or provided as raw data
- No URL Support: URLs are not supported directly
- No Detail Parameter: The
detailparameter is not supported - Token Costs: Approximately
(width × height) / 750tokens
Google Gemini
- Base64 Only: Images must be base64-encoded, URL references not supported
- No Detail Parameter: The
detailparameter is not supported - Token Usage: Tracks token usage in response metadata
Handle Incompatible Image Inputs
DSPy.rb raises ArgumentError or DSPy::LM::IncompatibleImageFeatureError for the incompatible model and input combinations below:
begin
# Attempt to use vision with non-vision model
non_vision_lm = DSPy::LM.new('openai/gpt-3.5-turbo', api_key: ENV['OPENAI_API_KEY'])
image = DSPy::Image.new(url: 'https://example.com/image.jpg')
non_vision_lm.raw_chat do |messages|
messages.user_with_image('What is this?', image)
end
rescue ArgumentError => e
puts "Error: #{e.message}" # Model does not support vision
end
begin
# Attempt to use URL with Anthropic (not supported)
anthropic_lm = DSPy::LM.new('anthropic/claude-4', api_key: ENV['ANTHROPIC_API_KEY'])
image = DSPy::Image.new(url: 'https://example.com/image.jpg')
anthropic_lm.raw_chat do |messages|
messages.user_with_image('What is this?', image)
end
rescue DSPy::LM::IncompatibleImageFeatureError => e
puts "Error: #{e.message}" # Anthropic doesn't support image URLs
end
begin
# Attempt to use URL with Gemini (not supported)
gemini_lm = DSPy::LM.new('gemini/gemini-2.5-flash', api_key: ENV['GEMINI_API_KEY'])
image = DSPy::Image.new(url: 'https://example.com/image.jpg')
gemini_lm.raw_chat do |messages|
messages.user_with_image('What is this?', image)
end
rescue DSPy::LM::IncompatibleImageFeatureError => e
puts "Error: #{e.message}" # Gemini doesn't support image URLs
end
Control Image Compatibility and Cost
- Choose a vision-capable model for image tasks.
- Resize large images to reduce token usage.
- Set image detail deliberately: use
lowfor simple queries andhighfor detailed analysis. - Check provider compatibility before sending images.
- Measure token use for multiple or large images.
Example: Object Detection
# Detect objects in an aerial image
airport_image = DSPy::Image.new(
url: 'https://example.com/aerial-airport.jpg'
)
response = lm.raw_chat do |messages|
messages.system(<<~PROMPT)
You are an object detection system.
Identify and count all airplanes in the image.
Provide approximate locations if possible.
PROMPT
messages.user_with_image('Detect airplanes', airport_image)
end
puts response
Token Usage Considerations
Images consume tokens based on their size:
- OpenAI: Varies by model and detail level
- Anthropic: Approximately
(width × height) / 750tokens
Monitor token usage when working with multiple or large images. raw_chat returns the accumulated text as a String; subscribe to lm.tokens for usage data (see Observability):
DSPy.events.subscribe('lm.tokens') do |_event_name, attributes|
puts "Tokens used: #{attributes[:total_tokens]}"
end
lm.raw_chat do |messages|
messages.user_with_image('Describe this', image)
end
Limitations
- File Types: Only JPEG, PNG, GIF, and WebP supported
- Size: Maximum 5MB per image
- Medical Images: Not suitable for medical diagnosis
- Text Recognition: May struggle with small or rotated text
- Spatial Reasoning: Limited precision for exact measurements
Run the Repository Examples
Complete working examples are available in the repository:
Bounding Box Detection
- File:
examples/multimodal/bounding_box_detection.rb - Features: Type-safe bounding boxes with
T::Struct, object detection, normalized coordinates - Use Cases: Aerial image analysis, object counting, computer vision tasks
Image Analysis
- File:
examples/multimodal/image_analysis.rb - Features: Color extraction, mood detection, and artistic analysis fields
- Use Cases: Art analysis, photography assessment, content moderation, image cataloging
Both examples include:
- Structured signatures with complex output types
- Type-safe multimodal processing
- Error handling and provider compatibility
- Integration with different vision models
Choose Provider and Output Types
- Run the examples locally with your API keys
- Define a signature for the output your application consumes
- Use rich types for nested structured outputs
- Check provider documentation for model-specific features