Inferencer v1.10.8 with MLX-INF

2026-03-27

In version 1.10.8, we added support for Sarvam MoE and Sarvam MLA, introduced MLX-INF for higher-accuracy quantization, improved distributed compute with persistent prompt caching, and made chat navigation faster with new keyboard shortcuts.

MLX-INF

We’ve added support for MLX-INF, a higher-accuracy quantization method designed to help preserve more model quality when running quantized models locally.

This is especially useful for users who want to balance memory usage with output quality, giving you more flexibility when choosing how a model is loaded and run on your device.

Sarvam MoE and Sarvam MLA

We’ve also added support for Sarvam MoE and Sarvam MLA, expanding the range of models you can run locally with Inferencer.

This brings more model coverage to the app and gives users additional options for local inference, experimentation, and private AI workflows.

Distributed Compute Improvements

Persistent prompt caching is now available for distributed compute.

This helps reduce repeated prompt processing across distributed inference sessions, which can improve responsiveness and efficiency when working with longer prompts, multi-turn chats, or agent-style workflows spread across multiple machines.

Faster Chat Navigation

We’ve added Command + Up and Command + Down for fast scrolling.

This makes it much easier to jump quickly to the top or bottom of a chat, especially during long conversations or when reviewing large amounts of generated output.

More Improvements

In this release, we’ve also:

  • Added fixes for DeepSeek 3.2 MLA
  • Added more bug fixes and performance improvements

As always, if you have any features or suggestions, you’re more than welcome to add them to our public roadmap.

Thanks again for your support, and as a reminder, all AI processing is done offline, directly on your devices. No telemetry, no background “update” checks.

P.S. If you find Inferencer useful, please consider leaving a review on the App Store. It would be much appreciated.

Inferencer

Artificial Intelligence should not be a black box.

Download Inferencer