One home for every local model on your machine
Every text, image, and speech model you already have, pulled through Ollama, cached by Hugging Face, or dropped in by hand, discovered and served from a single binary: a terminal UI to manage them, a local gateway to serve them, and coding harnesses seated on them in one command. Local-first, private, and free.
curl -fsSL https://hedos.ai/install | bashInstalls the hedos binary on macOS and Linux.
The whole shelf on one screen.
hedos shelf keeps the things a command line can’t keep in view: what is loaded and by whom, how much memory is left, what the gateway served today. Every key is a subcommand. The screen below is the shelf itself, drawn the way your terminal draws it: click it and use the keys it shows. The models’ replies and downloads here are simulated.
Pull, warm, unload, remove, chat, launch a harness, serve: the UI steps aside for anything that needs the terminal and is back the moment it ends. It works over ssh and inside tmux.
Found where they live. Run however they run best.
Hedos finds every model already on your machine, no matter who installed it, and serves each one on whatever engine runs it best. Discovered, then run.
A model that resolved a runtime is ready to serve. One that did not says so, instead of failing the moment you use it.
Install a model without leaving the shell.
Search Hugging Face by name or pick from what fits your memory. Hedos plans the install, shows you the size and the destination, and asks before a byte moves.
Downloads land in the standard Hugging Face cache or through the Ollama daemon, so every other tool on the machine sees the model too. Hedos owns no weights directory.
Seat a coding agent on a local model.
hedos launch runs Claude Code, OpenCode, Aider, Goose, or Crush against a model on your shelf. A gateway starts inside the same process on a free port, the harness is pointed at it, and both stop together. Your own harness config is never touched.
Before the harness starts, hedos sends one request shaped like the ones it will send, so a stopped daemon or a model without tool calling fails here with a reason, not inside the agent as a mystery.
One local endpoint for everything you own.
hedos serve speaks the OpenAI, Ollama, and Anthropic wire formats on loopback. Point an editor, an agent, or a script at it and reach every model on the shelf, tool calls included.
It binds to 127.0.0.1 and stays there. The only server here is the one on your desk.
Every sense the shelf has, one command each.
Text is where most of it happens, but the shelf holds vision, speech, transcription, and image models too, and each one has a verb that reads the same way.
$ hedos run llava "describe this" --image photo.png
Ask a vision model about an image. Named models are checked for sight before a byte moves.
$ hedos chat qwen3.5:9b
A conversation over stdin, streamed back, for scripts and pipes as much as for people.
$ hedos speak kokoro "good morning"
Synthesize speech to a WAV file through a local speech model.
$ hedos transcribe whisper voice.wav
Transcribe an audio file to text; the transcript streams as it is produced.
$ hedos image flux "a koala at a desk"
Generate a PNG through a local diffusion runtime, with progress on stderr.
$ hedos stats
Per-model usage from the gateway's audit log: requests, tokens, latency.
$ hedos warm gemma3 · hedos unload gemma3
Load a model ahead of the first request, or evict it, wherever it is held.
$ hedos ls --json
Every command speaks JSON on request, so the shelf is one jq away.