LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge
read at source ↗ huggingface.co
LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge
Source: HuggingFace Date: 2026-08-12 URL: https://huggingface.co/blog/LiquidAI/lfm2-5-vl-3b
Summary
Liquid AI’s LFM2.5-VL-3B pairs a 400M SigLIP2 vision encoder with its LFM2.5-2.6B text backbone (3.1B params total, pretrained on ~34T tokens), tuned for screen/UI comprehension, object grounding, and function calling. It answers directly rather than reasoning to stay fast, and reports 228 tok/s on M5 Max CPU and 20 tok/s with a ~3GB footprint on a Galaxy S26 Ultra, with day-one support across llama.cpp, MLX, vLLM, and ONNX.
Implications
Directly feeds local-model-feasibility: a 3B vision-language model with published mobile-phone throughput and memory numbers is a concrete edge-deployment data point for the tracked hardware tiers (M3 Max, M2 Max, WSL/3060), and the “answers directly, no reasoning” design choice is a legible trade-off between latency and capability worth comparing against reasoning-first small models.