Inferencer v1.9.1 for Coding Agents
In version 1.9.1, we added support for direct agent tool calling so you can use agent modes more reliably and directly from agentic development platforms such as VS Code.

Watch our agent comparison video here
P.S. If you find Inferencer useful, please consider leaving a review on the App Store. It would be much appreciated.
More Improvements
In this release, we’ve also:
- Added a detachable prompt progress indicator
- Added cache reuse for cancelled generations
- Improved model tool call detection for GPT-OSS, MiniMax and GLM
- Added support for Solar-Open, IQuestCode and EXAONE
- Added thinking support for EXAONE and Solar-Open
- Fixed Server API batching, thanks to alejandroed for the report
- Fixed messaging models that don’t support consecutive user messages - thanks to Fango2007 for the report
As always, if you have any features or suggestions you’re more than welcome to add them to our public roadmap.
Faster Inference for macOS 26?
In the next update currently in review by Apple, we’re adding specific improvements for macOS 26 to speed up the inference of prompt processing, but more on that soon.
Thanks again for your support, and as a reminder, all AI processing is done offline, directly on your devices. No telemetry, no background "update" checks.
