Inferencer v1.9.5 for Kimi-K2.5 and OpenClaw
In version 1.9.4 and 1.9.5, we added support for Kimi-K2.5 including thinking and distributed compute, as well as improved compatibility with OpenClaw and sped up the time to first token on macOS 26 by a whopping 32%.
P.S. If you find Inferencer useful, please consider leaving a review on the App Store. It would be much appreciated.
Kimi-K2.5
Last week we added support for Kimi-K2.5 a chart topping local AI model by the team at Moonshot.ai. In our tests with Inferencer, we achieved over 26 token/s running the 1 trillion parameter model on a single Mac Studio 512GB, with a batched inference speed of ~39 tokens/s across three inferences. Watch the test video here.
We also added support for toggleable thinking, custom expert control and distributed compute, where we sharded an even larger version of it across a Mac Studio and a MacBook Pro and utilised the expert control to answer a coding question that even ChatGPT and Claude failed on. Watch the test video here.
OpenClaw
Not just that, we added support for OpenClaw to allow for Kimi-K2.5 to run as the local AI agent.

Community
We also shared two compressed versions of Kimi-K2.5 on HuggingFace with both of them being in the top 5 most downloaded versions of the model this week.
As well as had another pull request merged into Apple’s MLX language model framework.
More Improvements
In this release, we’ve also:
- 32% faster time to first token (including load time) on macOS 26
- GLM-4.7-Flash: 29% performance bump and reduced memory usage
- Pro yearly plan
- Experts can be controlled via the server API using "experts": 10 or "experts": "faster"
- Batching support added for DeepSeek v3.2
- Server API and Tool call improvements
- Model streaming improvements for macOS 26
As always, if you have any features or suggestions you’re more than welcome to add them to our public roadmap.
More Coding Models
In the next update currently in review, we’re speeding up the inference of Qwen-3 Next models (including Qwen3-Coder-Next) by 28%, adding support for Step-3.5-Flash, amongst other things, but more on that soon.
Thanks again for your support, and as a reminder, all AI processing is done offline, directly on your devices. No telemetry, no background "update" checks.
