Inferencer v2.3.3 for Qwen 3.8 xHigh

2026-08-19

In version 2.3.3, we added support for Qwen 3.8 xHigh, Muse Glimmer and Ling-3.0-flash, along with improved loop detection, thinking detection, caching and inference performance across Qwen 3.8, Kimi K3 and other models.

Qwen 3.8 xHigh

Qwen 3.8 xHigh is now supported in Inferencer, and we’ve also added dedicated inference improvements for it. This should make Qwen 3.8 xHigh feel faster and more responsive during everyday use, especially when working with longer prompts or more demanding reasoning tasks.

Looping and Thinking

We’ve improved loop detection for Qwen 3.8 and various other models, helping reduce repetitive or stuck generations. We’ve also improved the thinking detector and added thinking limit support for MTP, giving you more control over how much thinking is used before a final answer is produced.

Caching and Performance

Caching improvements have been added for Qwen 3.8 and various other models, helping reduce repeated work and improve overall responsiveness. We’ve also improved inference for Kimi K3, alongside additional bug fixes and performance improvements across the app.

Improvements

  • Added support for Qwen 3.8 xHigh, Muse Glimmer and Ling-3.0-flash
  • Improved inference for Qwen 3.8 xHigh and Kimi K3
  • Improved loop detection for Qwen 3.8 and various models
  • Improved thinking detection and added thinking limit support for MTP
  • Improved caching for Qwen 3.8 and various models
  • More bug fixes and performance improvements

As always, if you have any features or suggestions you’re more than welcome to add them to our public roadmap.

Thanks again for your support, and as a reminder, all AI processing is done offline, directly on your devices. No telemetry, no background "update" checks.

P.S. If you find Inferencer useful, please consider leaving a review on the App Store. It would be much appreciated.

Inferencer

Artificial Intelligence should not be a black box.

Download Inferencer