Inferencer v2.2.0 for Files

2026-07-16

In version 2.2.0, we added file attachments and file tools, improved GLM 5.2 long context, added thinking limits to avoid endless thinking, and brought search, chat management, and downloading improvements to the app.

File Attachments and File Tools

You can now attach files directly to chats, including code, PDFs, Word documents, and ePub files. Inferencer can also perform text-only inference of images using metadata, OCR, and feature detection, making it easier to work with mixed document and image inputs without requiring a full vision model.

We also added file tools for local workflows, including Grep, Find and Replace, and Read/Write operations. Together with improved support for OpenCode, these changes make it easier to use Inferencer as a practical local assistant for coding, document review, and file-based tasks.

GLM 5.2 and Long Context Improvements

GLM 5.2 received a set of long context and inference improvements, including Extended EOG enabled by default. We also improved GLM 5.2 and MTP inferencing, helping longer and more complex prompts behave more reliably.

Alongside this, we added caching improvements for DeepSeek V4, Qwen 27B, and more, as well as broader caching improvements to reduce misses. Distributed compute also gained context precision support, and the Server API now supports reasoning, giving local deployments more flexibility for advanced workflows.

Thinking Limits and Better Model Support

Thinking can be powerful, but it can also loop or run for far too long. In this release, we added thinking limits to help avoid loops and reduce endless thinking more gracefully, while still allowing models to reason when needed.

We also improved thinking detection for more models, including Hy3, MiniMax M3, Ornith, and Step. Model support was expanded with Gemma 4 12B, MiniMax M3 tool calls, and Ornith-1.0 vision, giving you more options for local chat, coding, and multimodal workflows.

Search, Chat Management, and Downloading

We added Search in Conversation so you can quickly find previous messages without scrolling through long chats. Special tokens are now highlighted in the Entropy view, making it easier to inspect model behaviour and token-level details.

We also made chat management more flexible with new chats per window or tab, and you can now rename chats from the right click context menu. Downloading received improvements too, including concurrency, helping larger model downloads feel faster and more responsive.

More Improvements

In this release, we’ve also:

  • Added support for the Ollama API /ps endpoint
  • Improved support for OpenCode
  • Added engine stability fixes
  • Fixed an edge case crash and memory leak
  • Included additional bug fixes and performance improvements

As always, if you have any features or suggestions you’re more than welcome to add them to our public roadmap.

Thanks again for your support, and as a reminder, all AI processing is done offline, directly on your devices. No telemetry, no background "update" checks.

P.S. If you find Inferencer useful, please consider leaving a review on the App Store. It would be much appreciated.

Inferencer

Artificial Intelligence should not be a black box.

Download Inferencer