∫nferencer

Deeply control Artificial Intelligence models


Distributed inference
See the inner-workings
Fastest Inferencer
Markdown rendering
Serve models privately
More features

In version 2.3.7, we added support for DeepSeek V4 Vision and Hy4, expanded... Read More

Artificial intelligence
should not be a black box.

Private

By default, all AI processing happens offline and on your device. No data is sent to the cloud for processing.ⓘIf enabled, the server feature allows you to serve and connect to your own or trusted devices. No data is sent elsewhere.

State-of-the-Art

Support for SOTA models. Fastest inferencing performance. Patent Pending deep learning inferencing to fully control your models.

Realtime token probabilities

Token Inspection

Tap on the token inspector icon to reveal the probabilities for each token, showing you exactly what the model was thinking at each step.

Token Entropy: Instantly see contentious tokens, allowing you to better gaguge the confidence of the generation.

Token Selection: Specifically select tokens to manually explore alternative branches.

Token Exclusion: Select tokens to exclude from the generation, such as foreign characters.

Prompt prefilling

Expert Control

Unlock faster inference or increased intelligence by controlling the number of experts used in Mixture of Experts (MoE) models.

Prompt Prefilling

Expand the prompt field to seed a model's response by prefilling the Assistant message.

This technique allows you to direct its actions, skip preambles, structure outputs to enforce specific formats like JSON or XML, and even unlock gated responses.

Custom Tool Calls

The tools editor allows you to enable built in tools such as get_webpage_content or add in your own, so that the inferenced models can use them when needed.

Private hosted server

Private Server

Serve and inference models over the local network or internet with SSL encryption and IP security settings. Keeping the privacy of your inference in your premises.

Compatible APIs: Also includes Ollama and OpenAI compatible APIs for application development.

Mobile Support: Use Inferencer on iOS, iPadOS or visionOS to connect and inference larger models using your local compute.

Agent Support: Built-in tool calling support for agents such as GitHub Copilot, Continue.dev, Cline, Roo Code, Kilo Code, OpenClaw, OpenCode, Vibe CLI, Zed and more.

Persistent Prompt Caching: Caches processed prompts to disk, enabling sub-second response times for matched prefixes, even across distributed compute.

Distributed Inference: Link two or more computers together to run even larger models as a cluster.

Sandboxed Execution: Written in native code, Inferencer conforms to strict kernel-level sandboxing enforced by the operating system.

Private hosted server
Markdown rendering with LaTex example

More Features

Rendering: Support for markdown with advanced LaTeX rendering. Code previews also coming soon.

Batching: Continuously combines multiple requests to the same model into a single forward pass, dramatically improving total throughput (including support for batch caching for sub-second response times).

Model Streaming: Stream large models directly from storage, using custom read-only implementation for low-resource devices.

Model Compression: Optional compression during periods of inactivity to conserve memory for other applications, with a short delay on reuse instead of a full unload and reload.

Context Compression: Optionally reduces memory usage by quantizing the context window (including support for TurboQuant), allowing the model to handle larger contexts on low-memory devices.

Understand what the AI is thinking

Understand
why.

Latest Updates

Inferencer v2.3.7 with minMax Thinking

In version 2.3.7, we added support for DeepSeek V4 Vision and Hy4, expanded GLM-5.3 support, introduced new thinking modes, and improved tool calling and long-generation rendering.

Inferencer v2.3.3 for Qwen 3.8 xHigh

In version 2.3.3, we added support for Qwen 3.8 xHigh, Muse Glimmer and Ling-3.0-flash, along with improved looping, thinking detection, caching and inference performance across Qwen 3.8, Kimi K3 and other models.

Inferencer v2.3.1 with DSpark

In version 2.3.1, we added support for Kimi K3 and DeepSeek V4 DSpark, improved Kimi K3 inference, and added new controls to help stop thinking loops, manage timeouts, and set context limits.

Inferencer v2.2.2 for Thinking Loops

In version 2.2.2, we added thinking detection to help stop endless thinking loops, OpenClaw cache reuse optimisations, new model support, expanded tool calling, and a set of conversation and download improvements.

Inferencer v2.2.0 for Files

In version 2.2.0, we added file attachments and file tools, improved GLM 5.2 long context support, added thinking limits to avoid endless thinking, and brought search, chat management, and downloading improvements to the app.

Pricing

 
 

Base

Free

✓ Private on-device AI
✓ Sandboxed execution
✓ Unlimited processingⓘDefault response length is 2000 tokens
(Adjustable in Inference Settings)

✓ Markdown rendering
✓ Model control settings
✓ Download models
‒ Limited inference server
‒ Limited prompt cache
‒ Limited model streaming
✓ Model compression
✓ Context compression
‒ Limited distributed compute
‒ Limited batching
✓ Auto-load/unload settings
✓ Retention control settings
✓ Parental controls
‒ Limited custom tools
‒ Limited prompt prefilling
‒ Limited token entropy
‒ Limited token exclusions
‒ Limited token probabilities
‒ Limited expert control
‒ Limited Shortcuts
✓ Xcode Intelligence
✓ Visual Studio Code
 

Professional

$9.99 per month (USD1)
or $99.99 per year
✓ Private on-device AI
✓ Sandboxed executionⓘKernel-level operating system enforced application sandboxing
✓ Unlimited processing
✓ Markdown rendering
✓ Model control settings
✓ Download models
✓ Encrypted inference server
✓ Prompt CachingⓘEnables sub-second response times for previously processed prompts
✓ Model streamingⓘAllows streaming large models partially from storage for low-memory devices
✓ Model compressionⓘConfigurable on macOS 26.1 and above
✓ Context compression
✓ Distributed compute
✓ Multi-generation batchingⓘGenerations are merged together as one forward pass for higher total throughput
✓ Auto-load/unload settings
✓ Retention control settings
✓ Parental controls
✓ Unlimited custom tools
✓ Unlimited prompt prefilling
✓ Unlimited token entropy
✓ Unlimited token exclusions
✓ Unlimited token probabilities
✓ Mixture of experts control
✓ Shortcuts integration
✓ Xcode Intelligence
✓ Visual Studio CodeⓘIncluding support for agents such as GitHub Copilot and more
✓ Support the development
1. The prices shown here are in USD. Apple will convert the amount to your local currency at checkout, based on your region.

Subscribe for updates

With more features coming soon, you can be the first to know.