Inferencer v1.10.4 with Faster Model Loading
In version 1.10.4, we added support for Ling 2.5 and Ring 2.5, improved caching and batching behaviour, and sped up loading for standard Safetensors models by up to 6x.
Faster Model Loading
One of the biggest improvements in this release is a 6x faster loading path for standard Safetensors models. This should noticeably reduce the time between selecting a model and being able to start using it, especially for users who switch between models frequently or load larger checkpoints on macOS.
Faster loading also helps improve the overall feel of the app, making model switching, prompt regeneration, and background preparation less disruptive to your workflow.
New Models
We also added support for Ling 2.5 and Ring 2.5, along with chunked prompt support for Ling 2.5. Chunked prompting can help with larger prompts by making prompt processing more manageable and efficient.
In addition, we added support for loading community Qwen 3.5 models and fixed thinking support for Qwen 3.5. We also added thinking support for MiMo V2 Flash, giving you more control over how that model behaves during reasoning-heavy tasks.
Smarter Caching
Caching received a significant round of improvements in this release. We improved cache saving behaviour for prompt regeneration and added cache optimisations for a wide range of models, including GLM 4.7-Flash, GLM 5, GPT-OSS, Qwen 3.5, Llama 3.3, LongCat Flash Lite, MiMo V2 Flash, Qwen 3 Coder, Qwen Coder Next, Solar Open, Step 3.5, DeepSeek V3.2, Devstral 2 Small, Devstral 2, Mistral Small, MiniMax, and Kimi K2.5.
We also fixed next-turn cache usage when tool calls are made during batching, which should make multi-turn agentic workflows more consistent and efficient.
More Reliable Batching and Tool Calls
Batching and tool calling were another major focus in this release. We fixed a batching crash, resolved tool call issues during batching, and improved handling for tool calls in non-focused generations and background chats.
We also fixed control response handling for the new text splitting behaviour, which should reduce unexpected interruptions and make streamed responses behave more predictably.
More Improvements
In this release, we’ve also:
- Added model filtering based on system memory
- Improved download size reporting for Hugging Face
- Fixed Qwen 3.5 thinking support
- Added thinking support for MiMo V2 Flash
- Added support for loading community Qwen 3.5 models
- Fixed batching crash
- Fixed tool call issues during batching
- Fixed next-turn cache usage when tool calls are made during batching
- Fixed control response handling for new text splitting
- Fixed tool calls for non-focused generations
- Fixed background chat tool call handling
- More bug fixes and performance improvements
As always, if you have any features or suggestions you’re more than welcome to add them to our public roadmap.
Thanks again for your support, and as a reminder, all AI processing is done offline, directly on your devices. No telemetry, no background "update" checks.
P.S. If you find Inferencer useful, please consider leaving a review on the App Store. It would be much appreciated.
