Inferencer v1.9.3 for macOS 26

2026-01-29

In version 1.9.3, we added improvements for macOS including compressed memory support.

Memory Compression

By default, the model is now set to compress after 10 minutes of inactivity (togglable in the settings). This means that interactions with the model will be fast (like on macOS 15) but when it enters compressed mode, more RAM will be useable by other applications, without having to fully unload the model.

More Models

Also in this version we added support for GLM-4.7-Flash, LongCat-Flash-Thinking and Nemotron-3-Nano

More Models

Watch GLM-4.7-Flash in action

P.S. If you find Inferencer useful, please consider leaving a review on the App Store. It would be much appreciated.

More Improvements

In this release, we’ve also:

  • Added Korean support to DeepSeek v3.2
  • Improved local detection for the Server
  • Improved UI support for macOS 26
  • Fixed a crash with the Token Inspector

As always, if you have any features or suggestions you’re more than welcome to add them to our public roadmap.

Kimi K2.5?

In the next updates we’re adding even more performance improvements for macOS 26, distributed compute stability improvements and specific support for Kimi K2.5, but more on that soon.

Thanks again for your support, and as a reminder, all AI processing is done offline, directly on your devices. No telemetry, no background "update" checks.

Inferencer

Artificial Intelligence should not be a black box.

Download Inferencer