Image understanding with on-device AI
Apple's on-device Foundation Models framework gains multimodal image understanding and Vision Framework integration.
“This opens up new categories of experiences you can build with image understanding. It's as simple as attaching an image to your prompt.”
Apple announced multimodal capabilities for its on-device Foundation Models framework at WWDC, allowing developers to attach images to prompts for image understanding tasks. The update also integrates the Vision Framework, giving the model access to purpose-built tools like OCR and barcode scanning. This expands what developers can build with private, on-device AI without cloud dependencies.