Inferencer v1.11.6 with MTP Speculative Decoding

2026-05-26

In version 1.11.6, we added MTP speculative decoding for Qwen3.6 and Gemma4, persistent prompt caching for multimodal inference, multi-computer clustering, extended MCP support, and a number of performance and tooling improvements.

MTP Speculative Decoding

We added MTP speculative decoding for Qwen3.6 and Gemma4, delivering up to 2x faster inference depending on the model and workload. This is especially useful for long generations, agentic workflows, and repeated prompt processing where token generation can become the main bottleneck.

We also added Server API support for MTP, so the same speedups can be used from external clients, coding agents, and local automation workflows. For exact setup and expected gains, see the model description.

Persistent Prompt Caching

We added persistent prompt caching for multimodal inference, which can make cached requests up to 99x faster. This is particularly useful for workflows that repeatedly send the same images, documents, or long context blocks to the model.

The feature can be enabled in Settings, allowing you to keep the benefits of cached inference while still controlling how aggressively Inferencer stores and reuses prompt state.

Multi-Computer Clustering

We added multi-computer clustering for inferencing larger models across multiple machines. This makes it easier to run models that would otherwise be too large for a single system, while keeping the experience local and offline.

This is a major step forward for users who want to scale beyond a single Mac or workstation without moving to a cloud provider.

MCP and Tool Calling

We extended MCP support with additional integrations, including Playwright, MindsDB, Microsoft Learn, and DeepWiki. This makes it easier to connect Inferencer to external tools, knowledge sources, and automation workflows.

We also improved tool calling for DeepSeek V4, with an updated chat_template.jinja recommended for the best results. For Kimi K2.6, we improved both tool calling and inference behaviour, with an updated tokenizers_config.json required.

Thinking Levels

We added thinking levels for Ring 2.6, DeepSeek V4, OSS, and Hy3. This gives you more control over how much reasoning effort a model uses, allowing you to balance speed, cost, and answer quality depending on the task.

More Models and Performance

We added support for Mistral 3.5 Medium, expanding the range of high-quality local models available in Inferencer.

We also improved vision inference performance, including a 10% performance bump for Kimi K2.6 vision and a 16% performance bump for Qwen 3.5 and 3.6 vision.

More Improvements

In this release, we’ve also:

  • Added Server API support for MTP
  • Improved Markdown rendering
  • Fixed a vision inference cache bug
  • Added more bug fixes and performance improvements

As always, if you have any features or suggestions you’re more than welcome to add them to our public roadmap.

Thanks again for your support, and as a reminder, all AI processing is done offline, directly on your devices. No telemetry, no background "update" checks.

P.S. If you find Inferencer useful, please consider leaving a review on the App Store. It would be much appreciated.

Inferencer

Artificial Intelligence should not be a black box.

Download Inferencer