Inferencer 2.0.4 for GLM 5.2

2026-06-18

In version 2.0.4, we added support for GLM 5.2, brought multimodal inference to the Server API, improved image message handling, and made a number of improvements to the new multiprocessing engine, chat workflows, and macOS 27 beta compatibility.

GLM 5.2 and Model Validation

GLM 5.2 is now supported in Inferencer, giving you another strong local model option for chat, coding, and agentic workflows.

We also added validation for Kimi K2.7 Code and GLM 5.1 Code LoRA, helping ensure these models are handled more reliably across supported inference paths. DeepSeek INF models also now support MTP, expanding compatibility with newer DeepSeek-style model configurations.

Multimodal Server Inference

Multimodal inference is now available through the Server API, allowing image-aware requests to be handled directly by Inferencer’s local server.

This release also improves support for messages containing images, making it easier to work with vision-capable models through both the app and server-based integrations.

Server and Multiprocessing Improvements

We made further improvements to the new multiprocessing engine and caching for the Server API, along with improvements to multiprocessing queues. These changes should help make server-backed inference more stable and efficient, especially when handling repeated or concurrent requests.

We also added a Server API response length override, giving you more direct control over generated response limits when working with server-based clients.

Chat, Interface, and macOS 27 Beta

We added the ability to clone multiple chats, making it easier to branch conversations, compare outputs, or reuse context across different tasks.

New conversations are now faster to start, and we added a dark mode icon for a more consistent appearance in dark environments. We also added initial support for macOS 27 beta, helping prepare Inferencer for the next major macOS release.

Improvements

  • Added support for GLM 5.2
  • Added validation for Kimi K2.7 Code and GLM 5.1 Code LoRA
  • Added DeepSeek INF model support for MTP
  • Implemented multimodal inference via the Server API
  • Improved support for messages with images
  • Improved the new Multiprocessing engine and caching for the Server API
  • Improved multiprocessing queues
  • Added Server API response length override
  • Added support for cloning multiple chats
  • Made new conversations faster
  • Added a dark mode icon
  • Added initial support for macOS 27 beta
  • Added support for arc-1 SAP ADT MCP Tools Server
  • Included more bug fixes and performance improvements

As always, if you have any features or suggestions you’re more than welcome to add them to our public roadmap.

Thanks again for your support, and as a reminder, all AI processing is done offline, directly on your devices. No telemetry, no background "update" checks.

P.S. If you find Inferencer useful, please consider leaving a review on the App Store. It would be much appreciated.

Inferencer

Artificial Intelligence should not be a black box.

Download Inferencer