Liquid AI's DSpark Draft Models Boost LFM2.5 Inference Speed by Up to 3.2x
Liquid AI released DSpark draft model checkpoints for three models in its LFM2.5 family on August 20, 2026, introducing a speculative decoding path that boosts inference speeds by up to 3.2x on GPU and 2.87x on-device without altering output quality. For broader context, explore our AI News.
Understanding DSpark's Core Innovation
The DSpark draft models integrate a speculative decoding path that minimally increases memory footprint, adding approximately 300 million parameters per draft model. This design choice allows for a substantial speedup in decoding processes while maintaining the exact output quality of the baseline models. Liquid AI emphasizes that DSpark guarantees greedy-output parity, ensuring consistent results.
DSpark's architecture combines several key components: a DFlash-style parallel backbone, a lightweight sequential Markov head, and a confidence-scheduled verifier. This combination is engineered to optimize throughput and reduce latency in large language models.
Performance Gains Across Platforms
The integration of DSpark technology has led to considerable performance improvements. On GPU platforms, DSpark pushes throughput up to 3.18 times faster. For on-device applications, the speedup reaches up to 2.87 times. These figures highlight DSpark's capability to enhance the efficiency of AI models across different hardware environments.
Specifically, for the LFM2.5-2.6B model, DSpark reduces function-calling latency by an average of 57%. This particular model can achieve approximately 140 tokens per second on-device, demonstrating its potential for responsive local AI operations.
Model Availability and Ecosystem Support
The DSpark draft model checkpoints are available for three specific models in the LFM2.5 family: LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B. This release provides developers with immediate access to the enhanced capabilities.
Liquid AI has also ensured day-one support for popular inference frameworks, including llama.cpp and SGLang. This broad compatibility aims to facilitate easier integration and adoption of DSpark models within existing AI development workflows.
Implications for AI Development
The introduction of DSpark models by Liquid AI offers practical benefits for developers and organizations utilizing large language models. The significant increase in inference speed, coupled with guaranteed output quality, means that AI applications can run more efficiently, especially in scenarios requiring rapid responses or on resource-constrained devices. This could impact areas such as conversational AI, real-time data processing, and local AI deployments.
Conclusion
Liquid AI's release of DSpark draft models on August 20, 2026, represents a focused effort to improve the inference speed of its LFM2.5 family without compromising accuracy. By offering up to 3.2 times faster inference and broad framework support, DSpark aims to provide a more efficient foundation for various AI applications, particularly those requiring high throughput and low latency.
Sources
Recommended AI tools
ChatGPT
Conversational AI
AI research, productivity, and conversation—smarter thinking, deeper insights.
Perplexity
Search & Discovery
Clear answers from reliable sources, powered by AI.
Cursor
Code Assistance
The AI code editor that understands your entire codebase
DeepSeek
Conversational AI
Efficient open-weight AI models for advanced reasoning and research
Grok
Conversational AI
Your cosmic AI guide for real-time discovery and creation
GitHub Copilot
Code Assistance
Your AI pair programmer and autonomous coding agent
Was this article helpful?
Found outdated info or have suggestions? Send us a note.