Magnitude is not another chat window. It is an open-source inference engine for local agents that tunes kernels on your hardware before a model runs. The same folks who built an earlier browser agent (YC S25) just Launch-HN’d this as the thing they wished existed when they tried to run agents on-device.
Why local agents care
Agent sessions are long, bursty, and often concurrent. Datacenter batch engines optimize the wrong shape. Broad engines like llama.cpp chase compatibility. Magnitude’s pitch: compile/tune on the actual Metal, CUDA, Vulkan, or CPU box in front of you; hand-optimize popular open-weight families; grow and free KV memory as sessions start and stop; share prefix caches across sessions without tanking single-session decode.
Their published comparison against llama.cpp on Qwen 3.6 35B A3B (4-bit, 64k, no speculative decoding) claims about 92% faster decode on an M4 Pro and ~19% on a CUDA DGX Spark, with roughly 27% less memory per agent. Treat benchmarks as a starting point — HN commenters are already poking M5 gaps and MLX baselines — but the product shape matches how I actually run agents: several sessions, long context, machine still usable for other work.
Hook it to what you already use
It ships as a desktop app with one-click hooks for Pi, OpenCode, Hermes, Codex, and friends, plus an OpenAI-compatible endpoint. Apache 2.0. Nothing has to leave the box once the model is downloaded. On etherninja gear that is the interesting part — private inference next to the agent harness, not another SaaS meter.
They previously shipped a vision-first browser agent; the inference engine is the pivot after they got tired of local models being the weak link. Fair. Most of my agent pain lately is not “can it click,” it is “can I run two long sessions without melting the laptop or leaking prompts to a cloud meter.” Magnitude’s catalog trying to estimate fit and speed for your exact box is the right UX even if early M5 kernel gaps still need work.
Next step: install from magnitude.dev, let it assess your hardware, pull one model that fits, and point a coding agent at the local endpoint. See if decode and memory feel better than your current llama.cpp or MLX path.
Happy hacking.
References & Further Reading
- Magnitude – Product home — self-tuning local inference for agents.
- Magnitude Docs – Introduction – Features, backends, agent connections.
- Launch HN – Magnitude (YC S25) – Founders’ announcement and benchmark discussion.
- GitHub – magnitudedev/magnitude – Apache 2.0 source and issues.
- magnitude.run – Alternate product URL (same engine).
- AI TLDR – Magnitude – Independent summary of connects and API.