Inferencer v1.7.3 with Deep Speciale

2025-12-12

In version 1.7.3 we’ve added support for integrating Inferencer directly with GitHub Copilot, running DeepSeek V3.2 and advanced server management controls.

GitHub Copilot

GitHub Copilot support has been added to Inferencer, to set up, enable the Ollama compatibility API and use the Ollama API URL as the Copilot Ollama endpoint in VS Code settings. Watch this video for a demonstration.

GitHub Copilot

P.S. If you find Inferencer useful, please consider leaving a review on the App store, it would be much appreciated.

Would an Extremely Slow Einstein be Valuable?

Industry leaders Brian Case and Mark Cummings, PhD, featured Inferencer's model-streaming technology in Pipeline Magazine's top technology trends of 2025. Read more.

DeepSeek V3.2

Speaking of Einstein-level genius that can be run locally with model streaming on low memory devices… The big blue whale have released their chat topping V3.2 models, and Inferencer was one of the first AI engines to integrate support for them. Including releasing two of the top 10 most downloaded versions of their models to date.

DeepSeek V3.2

More Improvements

Also in this release you will find advanced security settings in the Server settings page, allowing you to specify API keys and block traffic using an IP allowlist or blocklist.

As always, if you have any features or suggestions you’re more than welcome to add them to our public roadmap.

What about Batching?

Also, if you want a sneak peek of Batching coming in the next version currently in review, watch this video.

Thanks again for your support, and as a reminder, all AI processing is done offline, directly on your devices. No telemetry, no background "update" checks.

Inferencer

Artificial Intelligence should not be a black box.

Download Inferencer