Inferencer v1.9 with Mixture of Expert Control

2026-01-11

In version 1.9.0 we added a new mixture of expert control, removed loading screens, moved distributed compute out of preview and more.

P.S. If you find Inferencer useful, please consider leaving a review on the App Store. It would be much appreciated.

Mixture of Experts

Using the new mixture of experts toggle, you are now able to control the number of experts utilised at inference time. This allows you to either speed up inference or potentially increase the intelligence. In our experiments, we managed to speed up the inference of GLM by almost 25% and still yield excellent code generation results, as well as improve MiniMax’s comprehension of riddles using this technique.

Mixture of Experts

Watch our experiments here

More Improvements

Also in this release, we’ve also:

  • Removed modal loading screens.
  • Moved distributed compute out of preview, so it’s now testable in the Free version (with limited tokens) on your systems.
  • Improved support for coding agents.
  • Reduced network traffic for client/server interactions.
  • Fixed a crash that was occurring on macOS 26 model downloads - thanks to Fango2007 for the report.

As always, if you have any features or suggestions you’re more than welcome to add them to our public roadmap.

Tool Calls for Coding Agents

In the next update currently in review by Apple, we’ll be adding tool calling for coding agents such as Continue, Cline, Roo Code and Kilo Code. But more on that soon.

Thanks again for your support, and as a reminder, all AI processing is done offline, directly on your devices. No telemetry, no background "update" checks.

Inferencer

Artificial Intelligence should not be a black box.

Download Inferencer