Inferencer v2.3.1 with DSpark
In version 2.3.1, we added support for Kimi K3 and DeepSeek V4 DSpark, improved Kimi K3 inference, and added more controls to help stop thinking loops, manage timeouts, and set context limits.
Kimi K3 and DeepSeek V4 DSpark
We added support for Kimi K3 and DeepSeek V4 DSpark, including multimodal model streaming from storage for Kimi K3. We also improved Kimi K3 inference and delivered a 17% performance boost to DeepSeek DSpark.
Thinking Detection
We improved thinking detection introduced in the last update, be sure to enable it in Controls to stop thinking loops.
Server API and Tool Calls
We fixed tool call handling and added Server API tool call omission. You can now disable Raw tokens in Server Settings to omit tool calls from the Server API when needed.
More Improvements
In this release, we’ve also:
- Improved Context Attention for DeepSeek and GLM
- Added inference timeout and context limit settings
- More bug fixes and performance improvements
As always, if you have any features or suggestions you’re more than welcome to add them to our public roadmap.
Thanks again for your support, and as a reminder, all AI processing is done offline, directly on your devices. No telemetry, no background "update" checks.
P.S. If you find Inferencer useful, please consider leaving a review on the App Store. It would be much appreciated.
