Inferencer v1.11.3 for MCP
In version 1.11.3, we added MCP Server support in Tools, faster model loads, new models, and improved inference for DeepSeek V4 and Kimi K2.6.
MCP Server Support
MCPCP Server support is now available in Tools, allowing Inferencer to connect to external tool servers and extend what local models can do without leaving the app. This makes it easier to bring structured tools, resources, and workflows into your local inference setup while keeping the core experience fast and private.
New Models
We added added support for Hy3, DeepSeek V4, MiMo V2.5 Pro and Ling 2.6, along with improved support for Kimi K2.6 and DeepSeek V3.2. These additions expand the range of local models available for coding, reasoning, and general chat, while giving you more options for balancing speed, quality, and memory usage.
Faster Loads and Inference
Model loads loads are now around 30% faster, reducing the time between selecting a model and getting started. We also improved inference for DeepSeek V4 and Kimi K2.6, with updated chat_template.jinja recommended for DeepSeek V4 and updated tokenizers_config.json required for Kimi K2.6. Caching and system prompt handling were also improved for vision models, helping multi-turn conversations stay more consistent and efficient.
Thinking Levels
Compatible models models can now use thinking levels, including DeepSeek V4, OSS and Hy3. This gives you more control over how much reasoning effort a model uses, allowing you to trade off speed and depth depending on the task.
Chat and Workflow Improvements
We added added message editing and conversation search, making it easier to refine prompts and find previous chats without scrolling through long histories. Loop detection was also improved, including better handling of code blocks, helping reduce repetitive output during longer coding and reasoning sessions.
More Improvements
In this release, we’ve also:
- Added MCP Server support in Tools
- Added support for Hy3, DeepSeek V4, MiMo V2.5 Pro and Ling 2.6
- Improved support for Kimi K2.6 and DeepSeek V3.2
- Improved inference for DeepSeek V4 and Kimi K2.6
- Added thinking levels for compatible models, including DeepSeek V4, OSS and Hy3
- Improved caching and system prompt handling for vision models
- Improved loop detection, including better handling of code blocks
- Added message editing and conversation search
- Made model loads around 30% faster
- More bug fixes and performance improvements
As always, if you have any features or suggestions you’re more than welcome to add them to our public roadmap.
Thanks again for your support, and as a reminder, all AI processing is done offline, directly on your devices. No telemetry, no background "update" checks.
P.S. If you find Inferencer useful, please consider leaving a review on the App Store. It would be much appreciated.
