Inferencer v1.8 with Continuous Batching
In version 1.8.0 we added a new batching engine for faster performance, contributed to MLX and added support for a new SOTA model.
Batching
To use, go to chat settings and enable the Batching option. Now all generations will be continuously piped through the new batching engine to be spliced together in the same model forward pass, which dramatically increases the total throughput.

Watch this video for a demonstration.
P.S. If you find Inferencer useful, please consider leaving a review on the App Store. It would be much appreciated.
Devstral-2 and MLX
As part of being one of the first apps to support Mistral’s new Devstral-2 model, we pushed out a fix for the underlying MLX engine, and we’re happy to report that our PR was accepted. So we’re officially part of the MLX codebase, as well as all the other apps that utilise it under the hood.

More Improvements
Also in this release, we’ve added the end of turn token to the token inspector, so you can force a conversation to continue, as well as fixes for edge case LaTeX rendering, multi-window sessions, and more.
As always, if you have any features or suggestions you’re more than welcome to add them to our public roadmap.
AI Media Generation?
We also launched our sister application xCreate which aims to let you generate AI media all locally and in-private on your device, starting with the state-of-the-art Z-Image-Turbo image generation model. For more information, visit xCreate.com or watch this video.
Thanks again for your support, and as a reminder, all AI processing is done offline, directly on your devices. No telemetry, no background "update" checks.
