Inferencer v1.10.2 with 28% Faster Qwen
In version 1.10.2, we added support for Qwen 3.5, Step 3.5, Devstral 2 Small and Mistral 3 Small. AS well as boosted the performance of Qwen 3 Next by 28%.
Thinking Improvements
We added support for Qwen 3 thinking, Devstral 2 Small thinking, and Mistral 3 Small thinking.
We also improved end-of-thinking handling for models that end with a think tag.
Distributed Compute and Stability
We fixed distributed compute issues for DeepSeek V3.2 and GLM 5, fixed an auto-load model crash related to a task already being launched, and added Transformers compatibility fixes.
More Improvements
In this release, we’ve also:
- Added a Clone conversation button
- Added Context MLA support
- Improved API tool call handling
- Added support for Qwen 3.5
- Added support for Step 3.5, including conversion, running, and batching support
- Added support for Qwen 3 Next experts
- Improved Qwen 3 Next performance by 28%
- Added support for Qwen 3 thinking
- Added thinking support for Devstral 2 Small and Mistral 3 Small
- Improved end-of-thinking handling for models ending with a
thinktag - Fixed distributed compute issues for DeepSeek V3.2 and GLM 5
- Fixed an auto-load model crash related to a task already being launched
- Added Transformers compatibility fixes
- Additional bug fixes and performance improvements
As always, if you have any features or suggestions you’re more than welcome to add them to our public roadmap.
Thanks again for your support, and as a reminder, all AI processing is done offline, directly on your devices. No telemetry, no background "update" checks.
P.S. If you find Inferencer useful, please consider leaving a review on the App Store. It would be much appreciated.
