Inferencer v1.9.3 for macOS 26
In version 1.9.3, we added improvements for macOS including compressed memory support.
Memory Compression
By default, the model is now set to compress after 10 minutes of inactivity (togglable in the settings). This means that interactions with the model will be fast (like on macOS 15) but when it enters compressed mode, more RAM will be useable by other applications, without having to fully unload the model.
More Models
Also in this version we added support for GLM-4.7-Flash, LongCat-Flash-Thinking and Nemotron-3-Nano

P.S. If you find Inferencer useful, please consider leaving a review on the App Store. It would be much appreciated.
More Improvements
In this release, we’ve also:
- Added Korean support to DeepSeek v3.2
- Improved local detection for the Server
- Improved UI support for macOS 26
- Fixed a crash with the Token Inspector
As always, if you have any features or suggestions you’re more than welcome to add them to our public roadmap.
Kimi K2.5?
In the next updates we’re adding even more performance improvements for macOS 26, distributed compute stability improvements and specific support for Kimi K2.5, but more on that soon.
Thanks again for your support, and as a reminder, all AI processing is done offline, directly on your devices. No telemetry, no background "update" checks.
