Magnitude launches a local inference engine tuned for AI agents
The open-source desktop engine tunes model kernels on local hardware and connects to existing agents; its speed claims come from the developer's benchmarks.
What launched
Magnitude’s developers introduced a local inference engine for AI agents in a September 30 launch post. The project is open source under Apache 2.0 and ships a desktop app for macOS, Windows and Linux. Its README says users can download an open model, then connect agents such as Pi, OpenCode or Codex; other clients can use an OpenAI-compatible API. The team describes support for Apple Silicon, NVIDIA and AMD GPUs, along with CPU-only machines. Compatibility and usable model size will still depend on the user’s hardware.
The technical pitch is device-specific tuning. Magnitude says it compiles and tunes kernels before a model runs, allocates memory as agent sessions grow and frees it when they stop. It also describes sharing cached context across concurrent sessions. These are the project’s stated design choices, not a guarantee of performance in every workload.
What the benchmarks do and do not show
In its launch post, the team compares Magnitude with llama.cpp using a four-bit Qwen model at a 64,000-token context. It reports faster decoding on an M4 Pro Mac and a DGX Spark, with smaller prefill improvements. The test conditions and numbers are supplied by Magnitude; no independent benchmark was reviewed for this story. The result may differ with another model, context length, quantization or device. The launch matters because it offers a downloadable, inspectable option for running agents locally, while its broad speed claim remains a claim to test.