Inferencer v1.3 with memory offloading

2025-10-08

Version 1.3 is now available for download on the Mac App Store. This update upgrades the token entropy inspector out of preview, and introduces a new memory offloading feature to allow for inferencing models larger than available memory.

Token entropy

The token entropy inspector allows you to instantly see contentious tokens - ones which were generated with low confidence.

Token entropy

Memory offloading

In preview is a new memory offload feature which allows you to partially stream models directly from storage. This can be incredibly slow depending on the size of the model you choose as storage typically have read speeds of 5GB/s, so if you’re fully streaming a 500GB model, expect just over 1 token a minute. But of course, for overnight inference operations, it can be incredibly useful to see how a larger or unquantized model performs. One more thing to note, is that this implementation of memory offloading, is completely read only, the data is only streamed from storage, not written to like typical implementations which can wear out solid state drives.

Memory offloading

More improvements

  • Support for Mistral
  • Model download and deletion improvements

Public roadmap

You can submit and vote features to be added to Inferencer in our public roadmap. In the next release we’ll be bringing inference queuing and shortcuts app integration, but more on that in v1.4.

If you have any suggestions you’d like to share in private, I’d love to also hear them.

Demo Video

If you’re interested to see how the GLM-4.6 performs, including the full unquantized version using memory offloading: https://youtu.be/bOfoCocOjfM

Thanks again for your support, and as a reminder, all AI processing is done offline, directly on your device. No telemetry, no background "update" checks.

P.S. If you find Inferencer useful, please consider leaving us a review on the App store, it would be much appreciated.

Inferencer

Artificial Intelligence should not be a black box.

Download Inferencer