Latest Updates
-
Inferencer v1.8 with Continuous Batching
In version 1.8.0 we added a new batching engine for faster performance, contributed to MLX and added support for a new SOTA model.
-
Inferencer v1.7.3 with Deep Speciale
The big blue whale have released their chat topping V3.2 models, and Inferencer was one of the first AI engines to integrate support for them. Including releasing two of the top 10 most downloaded versions of their models to date.
-
Pipeline Magazine December 2025
"Would an extremely slow Einstein be valuable?" - Industry leaders Brian Case and Mark Cummings, PhD, featured Inferencer's model-streaming technology in Pipeline's top technology trends of 2025.
-
Inferencer v1.7.2 with VS Code
Support for running local models directly through VS Code is now supported.
-
Inferencer v1.7 with custom tool calls
In version 1.7.0 the tools editor allows you to enable built-in tools such as ‘get_webpage_content' so that the models can use them when needed.
-
Inferencer v1.6.1 with Distributed ServerAPI
This update introduces support got multi-window sessions, distributed compute compatibility for OpenAI compatible API, code block copying and more.
-
Inferencer v1.6 with Distributed Compute
With Distributed compute you can pool the memory of two Macs together for inferencing larger models.
-
Inferencer v1.5.3 with Xcode Intelligence
Support for agentic code writing and interaction with Xcode Intelligence.
-
Inferencer v1.5 with Private Server
Private inference serving over the network and internet.
-
Inferencer v1.4 with Shortcuts.app integration
Generation queueing, shortcuts.app integration and faster startup times.
Subscribe for updates
With more features coming soon, you can be the first to know.