oMLX turns your Mac into a full LLM inference server, run from the menu bar. It serves text, vision, OCR, embedding and reranker models with continuous batching, plus a RAM+SSD tiered KV cache that survives restarts, so Claude Code and Cursor respond in about 5s instead of 90s. OpenAI and Anthropic compatible APIs drop straight in. Native Swift, not Electron. Apache 2.0, open source.
Open SourceDeveloper ToolsArtificial IntelligenceGitHub





Advertisement (Responsive)
💬 Comments
📋 Details
Launched
August 30, 2026
Upvotes
90
Comments
17
Source
Public data
Advertisement (Responsive)






This tool has 17 comments on Product Hunt.
View comments on Product Hunt →