Inferencer v1.6 with Distributed Compute
Version 1.6.0 is now available for download on the Mac App Store. This update introduces support for MiniMax M2, Kimi K2 Thinking and preview support for Distributed Compute.
Distributed Compute Preview
With Distributed compute you can now link together two Macs, sharing the memory together to inference larger models. The feature builds upon the model streaming and client/server architecture with specialised low latency networking to allow for inferencing extra large 500GB+ models such as:
- Qwen3-Code-480B at Q8 at 14 token/s
- Kimi-K2-Thinking Q4.25 at over 20 token/s
To use this preview feature, it must be enabled in both the app settings page and in the server settings. Once a connection to your server is made, if both the client and server have the same model, a distributed compute icon will appear next to the models dropdown list. Simply tap on it to load the models for distributed compute.

Please note that this is a preview feature, so if you run into any issues, please let us know. Once this feature is verified as stable, we’ll be adding in support for multiple computers. Stay tuned.
P.S. If you find Inferencer useful, please consider leaving a review on the App store, it would be much appreciated.
More improvements
- Support for MiniMax M2
- Support for Kimi K2 Thinking
- Model streaming support for Apertus
- Improved guardrails for model streaming
- Minor bug fixes.
Submit and vote on features to be added to Inferencer in our public roadmap.
Thanks again for your support, and as a reminder, all AI processing is done offline, directly on your device. No telemetry, no background "update" checks.
