Inferencer v2.3.7 with minMax Thinking

2026-09-04

In version 2.3.7, we added support for DeepSeek V4 Vision and Hy4, expanded GLM-5.3 support, introduced new thinking modes, and improved tool calling and long-generation rendering.

DeepSeek V4 Vision and GLM-5.3

DeepSeek V4 Vision and Hy4 are now supported, bringing stronger multimodal and model coverage to local inference. We also added support for GLM-5.3, GLM-5.3-Flash, Qwen3.8-Flash-Next, and Ling-3.0-Flash, giving you more options for fast, capable local models.

minMax Thinking

We improved thinking limit support for more models including GLM-5.3, along with additional thinking modes for Kimi K3. We also introduced minMax thinking, designed to deliver the strongest thinking behaviour while keeping token usage as low as possible.

Tool Calls and File Tools

Tool call streaming is now supported, making agent-style workflows feel more responsive and natural. We also fixed file tool behaviour for Grep and find-and-replace, improving reliability when working with codebases and local files.

More Improvements

In this release, we’ve also:

  • Fixed issue with DeepSeek MTP decoding
  • Improved rendering performance for long generations
  • Improved thinking limits and the end generation detector
  • More bug fixes and performance improvements

As always, if you have any features or suggestions you’re more than welcome to add them to our public roadmap.

Thanks again for your support, and as a reminder, all AI processing is done offline, directly on your devices. No telemetry, no background "update" checks.

P.S. If you find Inferencer useful, please consider leaving a review on the App Store. It would be much appreciated.

Inferencer

Artificial Intelligence should not be a black box.

Download Inferencer