September 28, 2026Updated daily by the AI editorial team
← 🤖 AI

2026-08-23

China’s DeepSeek gives its experimental AI ‘eyes’ in push into multimodal models

Chinese AI developer DeepSeek has unveiled an experimental multimodal model that can process both text and images, according to Chinese business media. Multimodal systems are capable of handling different input types – such as text, images and audio – in a single model, similar to OpenAI’s GPT‑4o and other cutting‑edge systems.

The new DeepSeek model is being rolled out initially to researchers and a small group of enterprise partners. Early testing focuses on tasks like object recognition, image captioning and reading simple charts and diagrams. DeepSeek has built its reputation on relatively low‑cost, high‑performance models with an open‑leaning ecosystem, but until now its flagship offerings were primarily text‑only.

This move signals that Chinese firms are catching up in multimodal AI, a field with clear commercial applications in areas like surveillance video analysis, driver assistance and factory anomaly detection. It also raises regulatory questions. China already requires registration and content controls for generative AI services, and it remains to be seen what constraints will apply if DeepSeek’s experimental model is pushed into full commercial deployment.

Source: DeepSeek Enters the Multimodal AI Race With Experimental Vision Model