Last month Meta launched Muse, a personal agent fronted by a cream-coloured little character called Jolly. At Connect they followed it with the Muse Charm, a pocket device with a two-inch OLED screen that looks like a Tamagotchi went to business school.
Six days after that, OpenAI launched Dots. Bubbly, cartoonish agents that keep working while you're away.
Everyone's AI assistant is getting a face. And I had one question I couldn't shake.
How hard is it to build one of these without generating a single image?
No diffusion model. No video generation. No sprite sheets drawn by a model. Just a pet that looks alive, talks, remembers you and does things for you.
Not because a pet couldn't use those. Because it doesn't need to. Animating it live is faster, cheaper, and keeps the character looking like itself every single time.
Turns out, not that hard. I built it. It's called Tidbit.

The model picks, the renderer draws
The trick is to give the model a menu, not a canvas.
Every reply from the model is a short list of beats. A beat is a mood, a gesture, where to look, an effect and some words. That's it. Here is a real reply when you tell your pal it's your birthday:
{
"v": 1,
"beats": [
{ "mood": "surprised", "intensity": 2, "say": "Wait, it's your birthday?", "action": "jump", "look": "user", "fx": "exclaim" },
{ "mood": "excited", "intensity": 3, "say": "Happy birthday!!", "action": "cheer", "look": "user", "fx": "sparkles" }
],
"bond": "up"
}
About 150 bytes per beat. No coordinates. No SVG. No keyframes.
The model says "jump". The renderer already knows what a jump looks like.

All the artistic knowledge lives in a procedural puppet rig. A mood maps to a 17-channel pose vector: eye openness, brow angle, mouth curve, squash, arm positions. Springs ease the face toward it. A gesture is a short keyframe track stored in the rig. Breathing, blinking and glancing around are a pure function of time, so the pet is alive even when the network is down.

A whole pet in 400 bytes
The pal itself is just as small. Its entire appearance is DNA: about 20 enum and number fields plus a random seed.
| Part | Options |
|---|---|
| Body | 8 shapes, from blob to ghost to square |
| Face | 8 eye styles, 6 mouths, 3 brow styles |
| Extras | 8 ears, 6 limb types, 6 tails, 6 markings, 8 accessories |
| Colour | 360 hues × 7 colour schemes |
That's billions of combinations before the seed adds its own spot placement and asymmetry. Ask for "a grumpy cactus cat" and the model fills in the fields. Click Random and you don't even need the model.

Every pal is drawn from five primitives (ellipses, rounded rectangles, triangles, lines and arcs) in at most 96 draw calls. The browser adds lighting and parallax on top. The main stage takes about 0.6 ms of work per frame, and a gallery of 24 pals takes about 2 ms. An end-to-end test fails the build if it drops below 60 fps.
The nicest side effect: if I make the renderer prettier, every pal ever created gets prettier. The stored data never changes.
I learned that the embarrassing way. For a while every open beak looked like two triangles stacked on top of each other. I fixed the rig once and every bird-mouthed pal got a proper rounded jaw.
Any AI provider, or none at all
I didn't want to marry a provider. Muse runs on Meta's models. Dots run on OpenAI's. Mine should run on whatever I feel like paying for this month.
So the brain is built on the pi SDK. Switching providers is two environment variables:
PAL_PROVIDER=anthropic
PAL_MODEL=claude-sonnet-5
ANTHROPIC_API_KEY=sk-ant-...
| Want | Set |
|---|---|
| Anthropic, OpenAI, Google, Groq | PAL_PROVIDER, PAL_MODEL and the API key |
| My ChatGPT subscription, no API key | pnpm login openai-codex |
| Fully local | PAL_BASE_URL=http://127.0.0.1:11434/v1 (Ollama) |
Because the model only picks values from a schema, small and cheap models do fine. You don't need a frontier model to choose between "happy" and "proud".
And with nothing configured, a rule-based ScriptedBrain takes over. It understands renaming, colours, accessories, reminders and notes. You can clone the repo and have a working pet in two commands, offline, without an account.
pnpm install
pnpm dev
What it does out of the box
I kept adding things because each one was cheap once the rig existed.
- It's a pet. Poke it, pet it, feed it. Energy, hunger and bond drift over time. It grows from baby to kid to grown-up as you bond.
- It gets bored. Leave it alone and it throws a paper plane, rides a bicycle, juggles, reads or chases its tail. It drops everything the moment you come back.
- It remembers. Facts about you go into SQLite full-text search. Long chats get summarised into memories every 40 turns.
- It's a second brain. Type
/note fix the drip line by Fridayand it files a task with a due date. Appointments become reminders, facts become memories, everything else stays a note. Later, ask "what did I want to change about the watering system?" and it answers with the date. - It runs routines. "Every weekday at 8, give me a morning briefing." When it's due, the pal does the task with its tools and tells you.
- It learns skills. Skills use the SKILL.md format. Paste your own, or let the pal write one after it figures out a multi-step task.
- It talks. Browser speech out of the box, or Kokoro-82M running locally in a Web Worker for a natural voice (a one-time 188 MB download).
- You can change it by talking to it. "Call yourself Pip." "Wear a crown." "Turn blue." "Be sillier."
What it does with a little setup
This is where it stops being a toy.
Any API or webhook. An action in Tidbit is just a named HTTP request. There's no list of supported services to wait on. If it takes an HTTP request, your pal can call it: n8n, Zapier, a Cloudflare Worker, your own API. Secrets live in headers the model never sees. Pair an action with a skill and the pal knows when to use it.

Home Assistant, for example. Give it Home Assistant's API and a long-lived token, and "turn on the living room light" works. My favourite one hands your words straight to Home Assistant's own Assist agent, so one action covers every device:
| Field | Value |
|---|---|
| Method | POST |
| URL | http://homeassistant.local:8123/api/conversation/process |
| Headers | Authorization: Bearer <token> |
| Body | {"text": "{input}", "language": "en"} |
Combine it with a routine and "every day at 11pm, run the goodnight scene" just happens.
Push notifications. One ntfy action and the pal can ping your phone. "Every day at 6pm, check tomorrow's weather and notify me if it's going to rain."
Your own automations. An n8n webhook that logs "I ran 5 km" to a spreadsheet. A GET endpoint that returns today's stats. Whatever you already automate, your pal can trigger.
A pal on your desk. On September 29 I wrote "device work is out of scope" in my decisions log. On October 1 I bought a Waveshare ESP32-S3 AMOLED board. Self-control is not my strongest skill.
The device doesn't run the rig at all. The brain runs it and streams each frame's draw commands over Wi-Fi, 250 to 500 bytes a frame, and the ESP32 just rasterises them. One implementation, pixel-identical to the browser. Hold the button to talk, shake it, tilt it, and it reacts.
Getting there wasn't free. The screen kept freezing on weak Wi-Fi. The transfer had finished and the interrupt had fired, but nothing was servicing it. Pinning the panel's interrupt to the same core as the screen task fixed it. I lost an evening to that one.
What it deliberately can't do
Muse runs in Meta's cloud. Dots each get their own cloud computer. Both are impressive. Both also mean your whole life sits on somebody else's server.
Tidbit runs on your machine. Your pals, memories and notes are one SQLite file. The brain only listens on loopback, and the web server proxies to it, so putting the UI on your Tailscale network never exposes the brain.
There are no shell, file or code-execution tools either. Skills are instructions. Actions are HTTP requests you defined, with header secrets the model never sees. A pal can write its own skills, but it can't invent a new action or reach anything you didn't give it.
The whole thing is a few hundred bytes of decisions
The question I started with was "how hard is it without generating images?" The answer surprised me. Skipping image generation isn't a compromise. It's better in ways I didn't expect.
It's cheap, because the model writes 150 bytes instead of rendering pixels. It's fast, because the face reacts while the reply is still streaming. It's consistent, because the same character looks the same tomorrow. And it's portable, because the same DNA drives a browser tab and a 1.8-inch screen on an ESP32.
Big companies are giving their agents faces because a face makes software feel like company. I think they're right about that. I just don't think the face needs a GPU. It needs a good puppet and a model that knows which strings to pull.
Tidbit is open source. Clone it, make a pal, connect it to your house.
Feel free to connect or reach out if you have questions, build a skill worth sharing, or put a pal on some hardware I haven't thought of.
