Inferencer v2.2.2 for Thinking Loops
In version 2.2.2, we added thinking detection to help stop endless thinking loops, OpenClaw cache reuse optimisations, new model support, expanded tool calling, and a set of conversation and download improvements.
Thinking Detection
Thinking can be extremely useful, but in some cases it can get stuck circling the same reasoning over and over. To handle this, we’ve added Thinking Detection, which you can enable in Controls to stop thinking loops and gracefully continue the response.
We’ve also added thinking limits to avoid loops and reduce endless thinking more gracefully. This gives you more control over when a model should keep reasoning and when it should move toward a final answer.
OpenClaw and Server API
We’ve added Server API cache reuse optimisations for OpenClaw, enabled by default, to make repeated agent interactions more efficient.
The Server API now also reports the full context window in context used, making it easier to understand how much room remains during longer sessions. We’ve also included excess prompt prefix fixes for Hy3 and other models, improving reliability for agent and API-driven workflows.
New Models and Tool Calls
This release adds support for Laguna S2.1, Inkling, Gemma 4 12B, and MiniMax M3 Tool Calls. We’ve also added Qwen 3.6 27B tool call support, expanding the range of models that can be used more effectively with agentic workflows.
GLM 5.2 also receives long context improvements, with Extended EOG enabled by default to help the model handle longer conversations and larger context windows more reliably.
Conversation and Interface
We’ve made it easier to work with long chats by adding Search in Conversation, so you can quickly find earlier messages without scrolling through everything.
We’ve also added special token highlighting in the Entropy view, new chats per window or tab, and the ability to rename chats from the right-click context menu. Copy and paste for images is now supported as well, making it easier to bring visual context into your chats.
More Improvements
In this release, we’ve also:
- Added downloading improvements, including concurrency
- Improved overall stability with additional bug fixes
- Continued performance improvements across the app
As always, if you have any features or suggestions you’re more than welcome to add them to our public roadmap.
Thanks again for your support, and as a reminder, all AI processing is done offline, directly on your devices. No telemetry, no background "update" checks.
P.S. If you find Inferencer useful, please consider leaving a review on the App Store. It would be much appreciated.
