Skip to main content
Mirobody can run with no API key: one model answers questions, another reads documents, both on the same machine (local-models.md is how to run them). This page records what was measured on 2026-09-30, what it means for photos, and the plan for the whole thing to fit on an ordinary computer.

Conclusion

  • Today’s default (1.5.4) is the small size: MiniCPM5-2B (Q4_K_M, 1.6 GB) answers and GLM-OCR-0.9B (1.4 GB) reads documents, 3.0 GB together on llama.cpp, in 16 GB of memory with no GPU. On the evaluation now published in benchmarks/local_models/ it passes 19 of 24 questions (Claude Code grade 215 of 248), stores all 140 printed rows of its 12 documents and writes 22 of 31 journal entries; before 1.5.4’s harness changes (below) it passed 16 of 24 questions, stored 45 of 140 printed rows and wrote none of 31 journal entries.
  • The large size is Qwen3.8-27B (GSQ-RCO IQ3_S, 13 GB), measured below: as correct as a hosted model on our questions, and it sees photos, but it needs about 20 GB of memory: a Mac with 32 GB, or a GPU with 24 GB.
  • 1.6.0 ships the goal: the same two small models post-trained for Mirobody, as Mirobody’s own model (Then, training), about 3 GB together, in 16 GB of memory and without a GPU, like the small size today. Two models, not one merged model. It replaces the default only when it passes the same evaluation as the model it replaces.
  • Photos: GLM-OCR reads printed text and tables, nothing else. Understanding what a photo shows, such as the calories on a plate or a rash, needs a model that sees: the large size, or a hosted one. The small pair does not, and Mirobody says so rather than guess.
  • Beside hosted models: model-choice.md puts the small size next to five cloud models (DeepSeek V4.1 Flash, Claude Sonnet 5.5, Claude Opus 5.5, Gemini 3.8 Flash and GPT-6 Luna) on the same cases, with what each costs and who reads the data.

What was measured

Apple M4 Pro with 48 GB, llama.cpp b11269, the demo record. Eight questions, each asked twice: a three-month trend chart, the latest value, change over time, medications, a lab report, a genotype, a record shared by someone else, and general knowledge. A run passes when the right tool is called with the right view, a chart is drawn when asked for, the answer is in the question’s language, and it answers. Then every number in every answer was checked against the database, and every “high” or “normal” against the range the report printed. Timings compare only within one session on one machine. The harness that produced these was never published; the evaluation that replaced it is benchmarks/local_models/: 24 questions, 12 documents and 15 journal sentences through the product’s own API, with cloud references run beside the local sizes.

MiniCPM5-2B, in detail

It is fast (most answers in 2 to 70 s) and gets the structure right, and its failures are narrow:
  • Asked on 2026-09-30 for “the past three months”, it passed no dates, received a year of real daily readings, and answered with October to December 2026: 92 points and three monthly means that do not exist.
  • One question (a shared record’s cholesterol) had no answer in either run (594 s and 375 s).
  • Asked for a reading as JSON, it put the value into the name.
MiniCPM5-1B, at Q8_0 so that quantization is not the excuse, called no tool in 12 of 16 runs and asked the user which “view” and time zone to use instead, so the 2B model is the training base. None of these is missing knowledge. They are dates, arithmetic and following the tool’s shape, which is what post-training fixes, and what the harness can take away from the model altogether.

Photos

GLM-OCR has four official prompts: Text Recognition:, Table Recognition:, Formula Recognition:, and information extraction into a JSON template. Four images (a monitor showing 128/91 and pulse 76, a Chinese nutrition table, the FDA sample label, and a plate of food): So Mirobody sends GLM-OCR only its text and table prompts (the local-ocr entry in config.llm.yaml), never a JSON template and never an open question, and the model that reasons over the text decides that 128 and 91 are a blood pressure. An estimate from a meal photo is a range at best, even from a model that sees. Since 1.5.4 the agent asks the model server whether its model can see (llama.cpp’s /props, Ollama’s /api/show; mirobody/utils/config/served.py). A model that cannot is sent a photo’s OCR text instead of the image, told that it is printed text only, and asked to say it cannot see the photo when the question is about the picture.

The plan

First, the harness

These help every model, hosted ones included, and leave a small model less to get wrong. Where 1.5.4 left each:
  1. Arithmetic in the tool. Means, monthly means, differences and trends come back computed; the model reports them. In place: view="stats" returns the count, minimum, maximum, mean, the first and last readings with their days and the change between them, and view="month" the monthly means; 1.5.4 made stats’ days the readings’ local days, and day, week and month answers say how a day is counted. What remains is the model using them: in the 2026-10-06 evaluation the small size averaged raw rows itself on two questions (p002-rhr-monthly, p003-weight-change).
  2. Dates in the tool. “The past three months” becomes a parameter the server resolves, so the model never computes a date. Remains.
  3. Charts cite the tool’s rows. The chart takes its points from the tool result instead of the model writing each one; the 92 invented points were written by hand. Remains: the model still writes the points. 1.5.4 only has the web client draw a chart with a stray brace, or say it could not.
  4. Every number checked before it is shown. A number in an answer that no tool result or document contains is flagged, and the answer regenerated. Remains: the evaluation’s score.py counts such numbers after the fact; the product does not check them.
  5. A shorter prompt for small models, measured the way the prompt in benchmarks/local_agent/ was. Remains for the agent. The journal’s request now opens with two worked answers; between the two evaluations MiniCPM5-2B’s journal went from 0 of 31 entries to 22 of 31 on the final code (24 after the first round of changes).
Also in 1.5.4, and not on the list (each in the CHANGELOG):
  • a table is read by its header with no model, whatever model is configured, from a born-digital PDF’s text layer, the printed page or an OCR model’s grid, and the text model gets only what the rules leave, told it is the rest of a medical report;
  • a long report is read a page at a time, and notes and logs give readings with their rows’ dates; a reading printed on two pages is stored once, and a page’s print date no longer dates its readings;
  • a looping extraction is bounded by its text and keeps its complete part;
  • a view asked for with no indicator answers with the catalogue, and minute to month views keep the newest 92 points instead of flooding the context;
  • a keyword finds a reading whatever its unit and spelling (FER, haemoglobin, a plural);
  • every JSON schema is closed, so OpenAI models read uploads and the journal;
  • a model’s two slots share one KV pool, so one question can use the whole context.

Then, training

This is what 1.6.0 ships, as Mirobody’s own model (model-choice.md). P0 is where 1.5.4 leaves it: the harness changes above; the evaluation published in benchmarks/local_models/ and benchmarks/local_ocr/, with MiniCPM5-2B’s baseline recorded (2026-10-06) and the cloud references beside it. Still to come: 200–500 questions split into train, dev and test. What is already in place:
  • Licences: MiniCPM5-2B is Apache-2.0 and GLM-OCR’s weights are MIT. Both publish fine-tuning routes (MiniCPM: TRL with PEFT, LLaMA-Factory, ms-swift, unsloth; GLM-OCR: a LLaMA-Factory guide).
  • A teacher: the large 27B model passes every run and judges against the printed range.
  • A reward that can be computed: whether each number in an answer is in the record is checked mechanically, which also filters the teacher’s runs.
  • Data at no labelling cost: the demo generator makes any number of people and series, and a report rendered from known values is labelled by construction.
Compute is modest: LoRA on a 2B model fits one 24–48 GB GPU, the 0.9B OCR model less. The default changes only when a trained model passes the same checks as the model it replaces.

Out of scope, for now

  • Understanding photos. MiniCPM5-2B is text-only and GLM-OCR reads text. With the small pair, a meal photo is answered with a request to describe the meal; photo understanding stays with a model that sees (the large 27B size, or a hosted key). A small vision model is a separate project.
  • One merged model. The two tasks need different data and different tests, and a regression in a merged model is hard to place. Mirobody already routes the agent and documents to separate entries, so two models drop in. A single small vision model doing both is worth measuring once both pass on their own.
  • Keeping it current. A change to a tool’s shape means another training run, so the evaluation runs nightly against the trained models.