llama-server serves them, and Mirobody
calls it. The shipped preset, the compose profiles, the first-run page’s
search for a server, the vision check (/props) and every measurement on this
page are llama.cpp’s. Any other OpenAI-compatible server can stand in for it
(Other servers), but llama.cpp is the one shipped and tested.
Two models, two jobs. GLM-OCR-0.9B reads report photos and pages into text
and tables; the tables’ rows are read by their column headers, with no model,
then coded by ② Translate, which is offline and deterministic. A second model
answers questions and writes titles, summaries and journal entries; it
comes in two sizes, below. Neither has to know a code.
All of them run on llama.cpp’s
llama-server, which has official builds for Windows, Linux and macOS (CPU,
Vulkan, CUDA, ROCm, Metal). One server serves them all:
docker/local-models.ini is a preset for its
router mode, and each model downloads from Hugging Face the first time it is
asked for.
Whether to run locally at all, and how the small size compares with the
cloud models on the same evaluation, is model-choice.md.
Choose a size
The setup page offers the same two, with these figures. Download is the GGUF files the preset fetches, document reader included; memory is the mostllama-server held with both models loaded while answering.
Small runs on any computer with 16 GB of memory and no GPU, Windows, Linux or
macOS; the stack beside it takes about 1 GB more. Large wants a 32 GB Mac or a
24 GB NVIDIA GPU. Two other answering models were measured and dropped:
MiniCPM5-1B answered 2 of the 24 questions with every expected fact for 0.4 GB
less download, and Qwen3.5-9B did not fit beside the stack on 16 GB.
The answering model is the only difference: both sizes read documents with
GLM-OCR and read tables by rule, so a lab report’s readings come out the same.
What changes is how well questions are answered, how fast, and whether a photo
in the chat is looked at (large) or read as its OCR text (small).
benchmarks/local_models/ has the
evaluation behind the figures, its cases and how to rerun it.
Without a GPU
The times above are Apple silicon’s, where llama.cpp runs on the GPU. On the CPU alone a first answer takes minutes. Measured on Linux on 2026-10-07, in llama.cpp’s CPU image (ghcr.io/ggml-org/llama.cpp:server, the local-cpu
profile below) with 4 vCPUs (colima, arm64):
- MiniCPM5-2B reads a prompt at about 50 tokens a second and writes at about 18. A 6.8k-token prompt was answered in 137 s, so expect 2–3 minutes for a first answer; later turns reuse the server’s prompt cache and read only what is new.
- GLM-OCR reads a photographed page in about 17 s, its two passes together.
- Both models loaded hold about 6.0 GiB in the container, so Docker’s VM
needs at least 8 GB of memory. Docker Desktop gives it half the computer’s
by default, 8 GB on a 16 GB machine; colima starts with 2 GB unless given
--memory 8. - The product’s tool-call probe (
doctor --probe) passed through this service, and both models load from its cache volume once downloaded.
Start the models
Pick the line for your machine. Every one serves the same preset; start it in themirobody folder (the one with deploy.sh). --models-max 2 keeps one
answering model and the reader in memory, so choosing another size on the
setup page unloads the one before.
Any computer with Docker, no GPU needed (Windows, Linux or macOS). The
slowest, a first answer in 2–3 minutes (Without a GPU), and
the one that needs nothing besides Docker, with at least 8 GB of memory for it:
deploy.sh writes COMPOSE_PROFILES into .env, so a later
docker compose up -d keeps the model service. On a stack that is already
running, docker compose --profile local-cpu up -d starts it beside the app
(the line the setup page shows; --profile local for the GPU service below);
add COMPOSE_PROFILES=local-cpu to .env as well, or a docker compose down
removes it and the next up leaves it out.
Windows, to use the GPU (Intel, AMD or NVIDIA, through Vulkan; the CPU when
there is none). winget’s ggml.llamacpp is llama.cpp’s own Vulkan release
build, for x64 and arm64. In PowerShell, from the checkout in WSL (below;
Ubuntu is the distribution’s name in wsl -l), so that the preset’s
relative path, the one the setup page shows, resolves:
host.docker.internal,
which on a Mac reaches a server on the host’s 127.0.0.1 (measured with colima).
This path has not been run on Windows by us; if the page finds no server there,
use the Docker line above, which needs no host networking.
Linux, or Windows with Docker Desktop, and an NVIDIA GPU, next to the app
(Linux needs the NVIDIA Container Toolkit):
llama-server line.
The first question after choosing a size waits for its download. Later starts
read the cache. Where huggingface.co is unreachable, point the download at a
mirror: HF_ENDPOINT=https://<mirror> in .env for the compose services, or
in the shell before llama-server. The models are the only thing fetched;
nothing about you is sent.
Mirobody itself on Windows runs in Docker Desktop (WSL 2 backend): open a
WSL terminal (Ubuntu), clone the repository there and run ./deploy.sh, as on
Linux. The scripts keep LF line endings on every checkout.
Point Mirobody at them
On the setup page (./deploy.sh prints its link; Settings → Model later),
choose 100% on this machine. Mirobody looks for the server at
host.docker.internal:8080, then compose’s llama:8080, then
127.0.0.1:8080, lists the models it serves, checks it serves the two you
choose (the preset’s, unless you pick others), asks it to load them, and
shows their progress. The choice is stored encrypted in the database.
Or two lines in .env, then docker compose up -d (a restart does not reread .env):
Other models
The answering model isminicpm5-2b unless you pick the large size
(qwen3.8-27b), and the document reader is glm-ocr: the sections of
docker/local-models.ini, which the local entries of config.llm.yaml ask
for. The setup page writes the size you pick; in .env the same is a line:
.env. A model chosen this way is checked the same way: the server
has to serve it. The candidates measured on the way to these two are in
local-models-roadmap.md. A model that cannot see is
detected from its server and sent a photo’s text instead of the photo.
The same works for a vendor key: the setup page shows the model beside the key
and takes another name (OPENROUTER_CHAT_MODEL=anthropic/claude-opus-5.5, for
one), checked with one real request before it is kept. Every entry’s variable
is its model_env in config.llm.yaml.
With a vendor key set as well, the key’s models come first, so the setup page
refuses local while .env holds a key. To use the local ones anyway, set
DEFAULT_MODEL=local, UTILS_VISION_MODEL=local-utils and
UTILS_TEXT_MODEL=local-utils.
compose.yaml maps host.docker.internal for Docker Engine on Linux and keeps
it out of HTTP_PROXY, so a proxied deployment does not route model requests
through the proxy. On a Mac (Docker Desktop or colima) that name reaches the
host’s loopback, so llama-server keeps its default 127.0.0.1. On Linux it
points at the Docker bridge, so a server on the host has to listen there: give
it the bridge address (--host 172.17.0.1, from ip -4 addr show docker0),
not 0.0.0.0, which also offers the server, with no key, to every machine on
your network. The local and local-cpu services need neither: the app
reaches them on compose’s own network.
failed to initialize router models: ... Is a directory in the llama log
means Docker could not see the checkout, and mounted an empty directory where
the preset should be. Colima shares only your home directory by default; keep
the checkout under it, or add the path to colima’s mounts.
The document reader
GLM-OCR-0.9B reads documents with either size. Three small OCR models that upstream llama.cpp serves were run through the product’s whole extraction path on synthetic reports (benchmarks/local_ocr/):
of 303 printed rows, GLM-OCR stored 283 with the printed value (302 with the
generator’s “SYNTHETIC SAMPLE” banner removed, as on a real report) and none
the page does not print; PaddleOCR-VL-1.6 278 (282), with 12 (8) the page does
not print; MinerU2.5-Pro 279 (299), with 3 (4). Of 301 handwritten rows they
stored 99, 62 and 55 (219, 136 and 171 without the banner). GLM-OCR stays the
default; model-choice.md
has the whole table and the reasons.
PaddleOCR-VL-1.6 is in the preset as [paddleocr-vl] (Apache-2.0, 1.8 GB
with its vision projector), and since 1.5.4 the product reads its answers’
OTSL tables and LaTeX units. Its prompts are not GLM-OCR’s, so
LOCAL_OCR_MODEL=paddleocr-vl alone is not enough: the local-ocr entry of
config.llm.yaml needs both changes:
mirobody and mirobody_worker in a
compose.override.yaml (model-choice.md
shows it), then docker compose up -d and doctor --probe.
Expect more rows in its OCR text, and on handwriting its passes looping to the
token cap (7 of 28 pages) and rows the page does not print.
Check it
doctor shows which entry each surface uses; --probe sends each one a real
request (a tool call, a schema-bound answer, a rendered image, the OCR passes)
through the code the product uses, and checks that each server runs the model
its entry names.
Other servers
Any OpenAI-compatible server works: set the two addresses and the model names it serves (curl <address>/v1/models), on the setup page or as LOCAL_MODEL
and LOCAL_OCR_MODEL.
Ollama. ollama pull qwen3.8:27b (17 GB) answered the large size’s earlier
set (8 questions, asked twice) correctly, and fastest. Two things to know:
- Ollama sets the context from the GPU’s memory, and below 24 GB it is 4,096
tokens, which cuts Mirobody’s prompt short without an error. Set
OLLAMA_CONTEXT_LENGTH=65536beforeollama serve. - Its
glm-ocrdoes not stop after reading a page in 0.34.1 and later (ollama/ollama#18609). Keep the documents on llama.cpp until that is fixed.
What each model can read in a photo
GLM-OCR reads printed text and tables. It cannot say what a photo shows: a meal, a rash or a scene is beyond it. Its official prompts are the only ones Mirobody sends (Text Recognition: and Table Recognition:, the local-ocr
entry’s ocr_prompts). Its JSON information-extraction prompt was measured
and is not used: on a blood-pressure display, a Chinese nutrition table and
an FDA label it put a value in the wrong field every time (systolic 76 for a
128/91 reading, energy “3” from the NRV% column), and asked for the dish on a
meal photo with no text it invented one.
What that means for each kind of photo (Apple M4 Pro, 2026-09-30):
The large size sees, so a photo in the chat reaches it as an image and a meal
can be estimated, roughly: expect a range, not a count. The default size does
not (MiniCPM5-2B; so does any model served without its
mmproj): that is
detected from its server, the agent gets the photo’s OCR text instead and is
told that is all it is, and for a meal it says it cannot see the photo and asks
what was eaten.
Things that behave differently from a hosted model
- Nothing streams while the prompt is read. A turn that adds 6.6k new
tokens waits over a minute for its first byte on an M4 Pro, over two on 4 CPU
cores. The
localentry allows 600 s of silence (stream_chunk_timeout); the library default of 120 s fails a long turn. - The first turn after loading is the slow one. The server caches the prompt it has read, so later turns read only what changed.
- A reply can be all reasoning. The agent asks once more when a reply has no answer text and no tool call; if the second one is empty too, the chat says it has no answer rather than showing a blank message.
How a document is read, whichever model reads the rest
These hold with a vendor key as much as here: the vendor’s model then extracts readings only from what the table rules left, though it still writes the file’s title and summary from the document’s text.- A long report is read a page at a time. Text over 3,000 characters with page headers goes to the model one page per request, two at once, and the pages’ readings are joined: in one request MiniCPM5-2B returned none of a 7-page check-up book’s 78 rows, page by page all of them. A log’s rows keep the dates they print.
- A table is read by its header. Rows under a header the rules know (项目名称 /
结果 / 参考值 / 单位, Analyte / Result / Unit, a CSV’s first line) are stored as
printed and labelled
rules:table@v1, when they look like readings and the page’s other copy (the text layer, or the OCR’s text pass) shows the same value. Patient details are skipped. The grid the OCR returns is read as the report printed it: a row whose empty cells moved is laid again by content, a header split over two rows or cells is joined, rows above a table’s first header borrow it, and a page with no header is typed by its cells when nothing about it is ambiguous. A row left unread, text outside the tables that holds a number or a finding, and a document with no table go to the text model, without the text pass’s copies of rows already read; a rule’s row outranks the model’s for the same printed row, under any name the vocabulary files under the same series. - Any OCR model’s answer is read the same way. An answer’s OTSL tables
(PaddleOCR-VL, MinerU) become HTML, its LaTeX (
\(\mu mol/L\)) the characters it typesets, and a line it repeats until the token cap one copy; each pass is capped at 8,192 tokens.