Inferencer v1.11.3 for MCP

2026-05-11

In version 1.11.3, we added MCP Server support in Tools, faster model loads, new models, and improved inference for DeepSeek V4 and Kimi K2.6.

MCP Server Support

MCPCP Server support is now available in Tools, allowing Inferencer to connect to external tool servers and extend what local models can do without leaving the app. This makes it easier to bring structured tools, resources, and workflows into your local inference setup while keeping the core experience fast and private.

New Models

We added added support for Hy3, DeepSeek V4, MiMo V2.5 Pro and Ling 2.6, along with improved support for Kimi K2.6 and DeepSeek V3.2. These additions expand the range of local models available for coding, reasoning, and general chat, while giving you more options for balancing speed, quality, and memory usage.

Faster Loads and Inference

Model loads loads are now around 30% faster, reducing the time between selecting a model and getting started. We also improved inference for DeepSeek V4 and Kimi K2.6, with updated chat_template.jinja recommended for DeepSeek V4 and updated tokenizers_config.json required for Kimi K2.6. Caching and system prompt handling were also improved for vision models, helping multi-turn conversations stay more consistent and efficient.

Thinking Levels

Compatible models models can now use thinking levels, including DeepSeek V4, OSS and Hy3. This gives you more control over how much reasoning effort a model uses, allowing you to trade off speed and depth depending on the task.

Chat and Workflow Improvements

We added added message editing and conversation search, making it easier to refine prompts and find previous chats without scrolling through long histories. Loop detection was also improved, including better handling of code blocks, helping reduce repetitive output during longer coding and reasoning sessions.

More Improvements

In this release, we’ve also:

  • Added MCP Server support in Tools
  • Added support for Hy3, DeepSeek V4, MiMo V2.5 Pro and Ling 2.6
  • Improved support for Kimi K2.6 and DeepSeek V3.2
  • Improved inference for DeepSeek V4 and Kimi K2.6
  • Added thinking levels for compatible models, including DeepSeek V4, OSS and Hy3
  • Improved caching and system prompt handling for vision models
  • Improved loop detection, including better handling of code blocks
  • Added message editing and conversation search
  • Made model loads around 30% faster
  • More bug fixes and performance improvements

As always, if you have any features or suggestions you’re more than welcome to add them to our public roadmap.

Thanks again for your support, and as a reminder, all AI processing is done offline, directly on your devices. No telemetry, no background "update" checks.

P.S. If you find Inferencer useful, please consider leaving a review on the App Store. It would be much appreciated.

Inferencer

Artificial Intelligence should not be a black box.

Download Inferencer