Reame

Reame

Self-hosted LLM inference on the hardware you already have

Reame is a CPU-first LLM inference server on llama.cpp with an OpenAI-compatible API. Built for narrow, repetitive workloads on cheap hardware — a €5 VPS, a free tier, a 2-core ARM box. Its memory layer caches prompts, prefixes and past generations to disk, so request #100 costs a fraction of request #1. MIT.

Open SourceDeveloper ToolsArtificial IntelligenceGitHub
👍 9 💬 5 comments 🕐 July 21, 2026
Advertisement (Responsive)

💬 Comments

This tool has 5 comments on Product Hunt.

View comments on Product Hunt →

📋 Details

Launched
July 21, 2026
Upvotes
9
Comments
5
Source
Public data
Advertisement (Responsive)

🔗 Related Tools