Latest Updates

  • Inferencer v1.11.1 with Multimodal Inference
    In version 1.11.1, we added full multimodal inference support, including image and audio models, along with model streaming support for Kimi K2.6, improved loop detection for code blocks, and a range of performance and stability improvements.

  • Inferencer v1.10.12 for ModelScope
    In version 1.10.12, we added ModelScope downloader support, macOS 26.2 M5 support, LongCat Next and Step 3.5 support, and a broad set of stability, UI, and Tasks view improvements.

  • Inferencer v1.10.8 with MLX-INF
    In version 1.10.8, we added support for Sarvam MoE and Sarvam MLA, introduced MLX-INF for higher-accuracy quantization, improved distributed compute with persistent prompt caching, and made chat navigation faster with new keyboard shortcuts.

  • Inferencer v1.10.4 with Faster Model Loading
    In version 1.10.4, we added support for Ling 2.5 and Ring 2.5, improved caching and batching behaviour, and sped up loading for standard Safetensors models by up to 6x.

  • Inferencer v1.10.2 with 28% Faster Qwen
    In version 1.10.2, we added support for Qwen 3.5, Step 3.5, Devstral 2 Small and Mistral 3 Small thinking, plus a 28% performance improvement for Qwen 3 Next.

  • Inferencer v1.9.5 for Kimi-K2.5 and OpenClaw
    Support for Kimi-K2.5 including thinking, distributed compute, expert control and OpenClaw. Plus, macOS 26 ttft speed up by a whopping 32%.

  • Inferencer v1.9.3 for macOS 26
    Improvements for macOS 26 and model support for GLM-4.7-Flash, Nemotron-3-Nemo, LongCat-Flash-Thinking and more.

  • Inferencer v1.9.1 for Coding Agents
    In version 1.9.1 we added support for agent tool calling via our Server APIs, allowing tools such as VS Code, GitHub Copilot, Continue.dev, Roo Code, Kilo Code and Cline to work better in agent mode.

  • Inferencer v1.9 with Mixture of Expert Control
    Recently we ran benchmarks of Inferencer against popular inferencing solutions only to discover some very interesting results.

  • Inferencer v1.8.2 with Non-Thinking
    In version 1.8.2 we added a new thinking toggle to the chat settings to allow you to turn off thinking for MiniMax M2.1 and GLM 4.7, along with added support for Server CORS, improved startup times and more

1 2 3 4

Subscribe for updates

With more features coming soon, you can be the first to know.