Inferencer v1.8.2 with Non-Thinking

2026-01-01

In version 1.8.2 we added a new thinking toggle to the chat settings to allow you to turn off thinking for MiniMax M2.1 and GLM 4.7, along with added support for Server CORS, improved startup times and more.

Faster Inference

Faster Inference

Recently we ran benchmarks of Inferencer against popular inferencing solutions only to discover some very interesting results.

P.S. If you find Inferencer useful, please consider leaving a review on the App Store. It would be much appreciated.

Non-Thinking

While thinking does have its benefits, especially for automated deeper and more correct answers, it can use up thousands of tokens circling over and over in some cases. So now you can turn off the preamble and provide the correct thoughts yourself. To use, go to chat settings and untick the Thinking option. The current supported models are MiniMax M2.1 and GLM 4.7, with more model support planned in the next updates.

Non-Thinking

Watch this video for a demonstration.

MiniMax M2.1 and MiMo-V2-Flash

Support for MiniMax M2.1 and MiMo-V2-Flash has also been added. These two models are notable for smaller activated parameters, meaning they inference very fast.

More Improvements

Also in this release, we’ve improved the app startup times for users who have hundreds if not thousands of chat messages. Remember the app supports an automatic deletion policy in the settings page which is worth exploring. We’ve also added context window size and memory to the stats page, added a CORS option to the Server Settings page to allow supporting web application access, and added support for a repetition penalty to Batched inferences.

As always, if you have any features or suggestions you’re more than welcome to add them to our public roadmap.

Master the Experts

In the next update currently in review by Apple, we’ll be adding the ability to control the number of experts used by MoE models to unlock faster performance or deeper intelligence. But more on that soon.

Thanks again for your support, and as a reminder, all AI processing is done offline, directly on your devices. No telemetry, no background "update" checks.

Inferencer

Artificial Intelligence should not be a black box.

Download Inferencer