Inferencer v2.3.1 with DSpark

2026-08-06

In version 2.3.1, we added support for Kimi K3 and DeepSeek V4 DSpark, improved Kimi K3 inference, and added more controls to help stop thinking loops, manage timeouts, and set context limits.

Kimi K3 and DeepSeek V4 DSpark

We added support for Kimi K3 and DeepSeek V4 DSpark, including multimodal model streaming from storage for Kimi K3. We also improved Kimi K3 inference and delivered a 17% performance boost to DeepSeek DSpark.

Thinking Detection

We improved thinking detection introduced in the last update, be sure to enable it in Controls to stop thinking loops.

Server API and Tool Calls

We fixed tool call handling and added Server API tool call omission. You can now disable Raw tokens in Server Settings to omit tool calls from the Server API when needed.

More Improvements

In this release, we’ve also:

  • Improved Context Attention for DeepSeek and GLM
  • Added inference timeout and context limit settings
  • More bug fixes and performance improvements

As always, if you have any features or suggestions you’re more than welcome to add them to our public roadmap.

Thanks again for your support, and as a reminder, all AI processing is done offline, directly on your devices. No telemetry, no background "update" checks.

P.S. If you find Inferencer useful, please consider leaving a review on the App Store. It would be much appreciated.

Inferencer

Artificial Intelligence should not be a black box.

Download Inferencer