Latest Updates
-
Inferencer v2.3.7 with minMax Thinking
In version 2.3.7, we added support for DeepSeek V4 Vision and Hy4, expanded GLM-5.3 support, introduced new thinking modes, and improved tool calling and long-generation rendering.
-
Inferencer v2.3.3 for Qwen 3.8 xHigh
In version 2.3.3, we added support for Qwen 3.8 xHigh, Muse Glimmer and Ling-3.0-flash, along with improved looping, thinking detection, caching and inference performance across Qwen 3.8, Kimi K3 and other models.
-
Inferencer v2.3.1 with DSpark
In version 2.3.1, we added support for Kimi K3 and DeepSeek V4 DSpark, improved Kimi K3 inference, and added new controls to help stop thinking loops, manage timeouts, and set context limits.
-
Inferencer v2.2.2 for Thinking Loops
In version 2.2.2, we added thinking detection to help stop endless thinking loops, OpenClaw cache reuse optimisations, new model support, expanded tool calling, and a set of conversation and download improvements.
-
Inferencer v2.2.0 for Files
In version 2.2.0, we added file attachments and file tools, improved GLM 5.2 long context support, added thinking limits to avoid endless thinking, and brought search, chat management, and downloading improvements to the app.
-
Inferencer v2.0.6 for MiniMax M3
In version 2.0.6, we added support for GLM 5.2, GLM 5.2 MTP and MiniMax M3, as well as Server API image support for OpenClaw and Cherry, drag and drop conversation reordering, and more.
-
Inferencer 2.0.4 for GLM 5.2
In version 2.0.4, we added support for GLM 5.2, brought multimodal inference to the Server API, improved image message handling, and made a number of improvements to the new multiprocessing engine, chat workflows, and macOS 27 beta compatibility.
-
Inferencer v2.0.1 with Multiprocessing
In version 2.0.1, we added a multiprocessing engine for running multiple models at the same time, expanded multimodal model support, and brought major performance improvements to DeepSeek V4.
-
Inferencer v1.11.6 with MTP Speculative Decoding
In version 1.11.6, we added MTP speculative decoding for Qwen3.6 and Gemma4, persistent prompt caching for multimodal inference, multi-computer clustering, extended MCP support, and a number of performance and tooling improvements.
-
Inferencer v1.11.3 for MCP
In version 1.11.3, we added MCP Server support in Tools, faster model loads, new models, and improved inference for DeepSeek V4 and Kimi K2.6.
Subscribe for updates
With more features coming soon, you can be the first to know.