Liquid AI ships a 2.6B on-device agent model
Liquid AI shipped LFM2.5-2.6B, an open-weight agentic model designed to run fully on-device. It is positioned for private local agent workflows, with reported speeds of 220 tokens per second on Apple M5 Max, 113 tokens per second on AMD Ryzen, and around 30 tokens per second on phones. Liquid says it beats Qwen3.5-9B on tool-use benchmarks despite being much smaller, though larger models still lead on coding.
- The model sits inside Liquid’s broader LFM lineup, where text models are aimed at tool calling and structured output alongside chat and classification use cases.
- Liquid’s documentation lists 32K token context as a shared capability across most models, with a larger 128K window reserved for LFM2.5-8B-A1B.
- Deployment options include local formats such as GGUF, MLX, and ONNX, giving developers multiple paths for CPU, Apple Silicon, edge, and production runtimes.
