orbit
Hackable always-on robots
Hackable always-on robots: one brain, many bodies. Sees, hears, talks, moves. Vision, audio and language models are all swappable, on hardware you extend yourself.
Every body streams what it senses into one Brain. Agents reason over it and act back out, and the robot senses its own actions, which closes the loop.
Today the body is a desk companion: an Android phone for the face, camera, mic and voice, on a 3D-printed base with servos driven by an ESP32. A camera-and-mic rover board works as a proof of concept, with no wheels yet.

What it does
- One brain, many bodies: each device streams what it senses to one Brain, where an agent reasons over it and acts back out.
- A desk companion you can build. The phone is the face, camera, mic and voice; an ESP32 drives the servos. The repo has OpenSCAD source and STL files for a planned body, a first pass not yet ready to print.
- Swappable models. Speech-to-text runs locally (Parakeet or Whisper). Vision runs on Gemini, or on a local Qwen2.5-VL on Apple Silicon. The language model is Gemini, or another model through OpenRouter.
- Perception and memory. Vision, speech, identity, emotion and body-motion snapshots are fused into a summary the agent reads each turn, and a per-bot store keeps the facts it chose to remember.
- Over-the-air updates for the phone app and the ESP32 firmware, served by the Brain. The body rolls back if a new build fails to rejoin.
- A browser console with a live roster, Perception Studio, body controls and traces, and one script that takes a fresh Ubuntu 24.04 box to a tested, running Brain.
What it doesn't do
- No app store. The phone app and the ESP32 firmware update from your own Brain.
- No API keys on the phone. Model, persona and provider keys stay in the Brain's config.
- No device-to-device links. The phone and the servo body each talk only to the Brain.
- No transcripts in long-term memory: it keeps discrete facts, not a dump of everything the robot perceived.
Good to know
- It is not local-only by default: the language model, vision and the perception summaries call Google's Gemini API with your own key. Speech-to-text runs on the Brain's machine.
- By default the Brain keeps conversation transcripts, recent audio clips and camera frames, six hours of raw perception and 30 days of hourly summaries on its own disk.
- The Brain's port has no login yet and no TLS by default; an IP allowlist is the only protection, so keep it on a network you trust.
Get it
The Brain needs Node.js 22 or later, Python for its speech-to-text sidecar and a Gemini API key: cd brain && npm install && npm run dev serves it and its console on port 8099, or ops/orbit.sh bootstrap sets up a fresh Ubuntu 24.04 box. The phone app is sideloaded from source onto Android 8.0 or later, and the ESP32 body is flashed once with PlatformIO; after that, both update over the air.
github.com/seeknull/orbit How the Brain works Run a Brain Body hardware notes