Ferrum is an open-source, Rust-native runtime for running and serving local language models on Apple Silicon Metal and NVIDIA CUDA. One binary gives you an interactive CLI plus OpenAI-compatible Chat Completions and Responses APIs, without requiring Python, PyTorch, or vLLM at runtime. Choose a model explicitly, run it locally, or expose the same model over HTTP.
Open SourceDeveloper ToolsArtificial IntelligenceGitHub



Advertisement (Responsive)
💬 Comments
📋 Details
Launched
September 2, 2026
Upvotes
3
Comments
1
Source
Public data
Advertisement (Responsive)






This tool has 1 comments on Product Hunt.
View comments on Product Hunt →