Inferencer v2.4.1
Inferencer for Claude Code, Instant Thinking and Multiserver Inference
In version 2.4.1, we added an Answer Now button to dynamically end thinking, expanded support for Step5 Preview and MiMo V2.6, and made it possible to inference multiple servers at the same time. We also improved tool calls, caching, batching, multiprocessing, and server switching to make Inferencer faster and more reliable for agentic workflows.
Answer Now and Thinking Control
The new Answer Now button lets you dynamically end thinking when you already have enough reasoning and want the model to respond immediately. This gives you more direct control over long-running thinking sessions, especially when a model is overthinking or when you want a faster final answer.

We also added thinking limit improvements, thinking limits for the Server API, and MTP support for thinking limits. For users who want stricter control, the thinking limit enforcer option is available but disabled by default.
Claude Code and Anthropic API Support
We added Anthropic API and Claude Code support, including cache optimisations, making it easier to use Inferencer with agentic coding workflows. This complements our existing tool calling improvements and helps local models integrate more smoothly with developer tools.
We also added OpenCode support for Server API date override, improved in-app tool calls, and made tool calls run client side when inferenced via the Server. These changes improve compatibility with coding agents and make tool-based workflows more predictable.
Multi-Server Inference and Server Switching
Inferencer can now inference multiple servers at the same time, giving you more flexibility when running distributed or multi-machine inference setups. We also integrated server switching into the models toolbar, so you can move between servers more quickly without leaving the main workflow.
This builds on our recent multi-computer clustering support, which allows larger models to be inferenced across multiple machines.
More Improvements
In this release, we’ve also:
- Added support for Step5 Preview and MiMo V2.6
- Added a Python code viewer
- Improved missed tool call detection and handling
- Improved Markdown code blocks
- Hardened batching and multiprocessing
- Improved caching for Qwen 3.8 and similar models
- Improved HTML import errors
- Added bug fixes and performance improvements
As always, if you have any features or suggestions you’re more than welcome to add them to our public roadmap.
Thanks again for your support, and as a reminder, all AI processing is done offline, directly on your devices. No telemetry, no background "update" checks.
P.S. If you find Inferencer useful, please consider leaving a review on the App Store. It would be much appreciated.

