2026-09-27
Liquid AI Debuts LFM2.5‑VL‑DSpark Draft Model to Turbo‑Charge Vision‑Language Inference
Startup Liquid AI has announced LFM2.5‑VL‑DSpark, a new “draft” model designed to accelerate its existing 3‑billion‑parameter vision‑language model, LFM2.5‑VL‑3B. DSpark is built around speculative decoding: a smaller model quickly proposes batches of likely next tokens, which the larger model then verifies in fewer, more efficient steps. The technique aims to cut latency and GPU time without changing the final output quality.
Unlike text‑only draft models, DSpark can handle both images and text, making it suitable for tasks like visual question answering and chat about screenshots or documents. When paired with LFM2.5‑VL‑3B, Liquid AI claims significant inference‑time speed‑ups, opening the door to cheaper deployment on modest hardware instead of top‑tier data‑center GPUs.
The company is releasing technical details and benchmarks for researchers and developers, positioning DSpark as part of a broader move to make multimodal AI more affordable to run. As cloud costs and power consumption become central constraints on large‑scale AI, optimizations at the inference layer—such as draft models, quantization, and smarter scheduling—are likely to matter as much as headline model size in determining who can compete in generative services and on‑device intelligence.