Overview
This module provides comprehensive health data file processing capabilities, including file upload, health indicator extraction, and file deletion. The system uses a modular design, supports multiple file types, and provides real-time progress feedback.Table of Contents
- Features
- Supported File Types
- File Upload
- Processing Flow
- Health Indicator Extraction
- File Deletion
- API Reference
- Data Models
- Error Handling
- Best Practices
- Code Examples
Features
- ✅ Multiple Upload Methods: Supports WebSocket real-time upload and REST API batch upload
- ✅ Real-time Progress Feedback: WebSocket connections provide real-time progress updates for file upload and processing
- ✅ Smart File Recognition: Automatically identifies file types and selects the appropriate handler
- ✅ Health Indicator Extraction: Uses LLM to automatically extract health indicator data from medical reports
- ✅ Multi-format Support: Supports PDF, images, Office documents, text, genetic data and more
- ✅ PDF Parallel Processing: Multi-page PDFs are processed in parallel for improved efficiency
- ✅ File Summary Generation: Automatically generates file content summaries
- ✅ Cascade Deletion: Automatically cleans up associated health data when files are deleted
Supported File Types
Two lists, and they are not the same list.
mirobody/utils/file_types.py is
the UPLOAD accept-list: what a user is allowed to hand the server. It carries
.xls and .xlsb, which no parser here reads. What a file can be READ as is
mirobody/documents/detect.py: EXTRACTABLE_SUFFIXES (15, through a parser
or OCR) plus TEXT_EXTENSIONS (8, decoded directly) is the 23 the README
counts. When this table and that module disagree, the module is the contract.
Whatever the handler, the TEXT of a document comes from one place:
mirobody/documents/ (detect.kind by extension, content type and, when
those lie, the bytes; extract.extract_text by kind). A PDF gives up its
embedded text layer page by page and only the pages that have none — scans —
are rendered and handed to the vision model, one image at a time; a photo is
downscaled and OCR’d; a spreadsheet or Word file never reaches a model at
all. The result is cached by content hash through th_files, so the same
bytes are never OCR’d twice, and the file summary is generated from that text
rather than from the file.
File Upload
WebSocket Upload (Recommended)
WebSocket upload provides real-time progress feedback, ideal for large file uploads and scenarios requiring real-time status updates.Endpoint
Connection Flow
upload_completed and extraction_completed answer different questions, and a
client needs both: the first says the file arrived and is readable, the second
says whether anything could be extracted from it. Extraction runs after the
upload has already completed, so a client that stops listening at
upload_completed will render a stored file as a finished one.
Message Types
Client sends:
Server response:
date_source is document when the date was read off the report itself, and
upload when none was found and the upload time is standing in as a
placeholder. A client should offer to correct the second case.
This is the event that says whether extraction worked, and the only one
that does. upload_completed above reports that the FILE arrived and is
stored — it says 1 files successful even when every model call failed, because
the file is genuinely stored and readable either way.
failed: true means extraction could not run — no provider is configured, or
every call to the configured one errored. The sentence explaining which is on
the file row (error, from GET /api/v1/health-indicators/files); the file
itself is stored and readable. failed: false with indicators_count: 0 is a
different and ordinary outcome: a document that genuinely has no indicators in
it. A client must not render the first two like the third — that is what made a
green row over an empty list, with the cause visible only in server logs.
Client sends:
Timeout Mechanism
REST API Upload
REST API provides a simple file upload method, suitable for simple scenarios or applications that don’t require real-time progress.Endpoint
Request Parameters
Response Example
Response Codes
Processing Flow
Architecture Overview
Processing Steps
- File Upload Phase (0-30%)
- Receive file data
- Validate file type and size
- Generate unique file identifier
- Upload to object storage (S3/OSS)
- File Type Recognition (30-35%)
The system automatically identifies file types via FileHandlerFactory:
- Content Processing Phase (35-90%)
PDF File Processing:
- Summary Generation Phase (90-95%)
- Generate file content summary using LLM
- Generate intelligent file name
- Result Saving Phase (95-100%)
- Save processing results to database
- Write the extracted indicators as observations (
th_observation, coded on the way in) - Update user health profile
Health Indicator Extraction
Description
The system uses Large Language Models (LLM) to automatically identify and extract health indicator data from medical reports.Configuration
Supported Indicator Types
- Complete Blood Count (WBC, RBC, Hemoglobin, etc.)
- Biochemistry (Liver function, Kidney function, Lipids, etc.)
- Physical Examination (Blood pressure, Heart rate, Weight, etc.)
- Tumor Markers
- Thyroid Function
- Other Medical Test Indicators
Extraction Result Format
Report Date
content_info.date_time becomes the start_time/end_time of every reading
extracted from the file. When the document shows no date (or one the parser
cannot read), the readings are filed under the user’s current time — and that
fallback is labelled, not silent: each reading’s comment JSON and the
file row carry date_source:
The labelling exists because extraction runs one file at a time. A report
photographed as three screenshots shows its date on the first page only, and
“page 2 of the same report” is indistinguishable from “a second report whose
date did not come out” — so nothing is inherited automatically. The files
listing (
GET /api/v1/data/uploaded-files) exposes report_date,
date_source and date_confirmed per file; the web client asks about
upload_time files and offers the dates read from the other files of the same
upload (same created_source_id, or uploaded within a few minutes of it — the
web client opens one upload session per file).
upload_completed:
report_date_detected (file_key, report_date, date_source, seconds after
the upload) and extraction_completed (adds indicators_count); both carry
the upload’s sessionId, which is how the web client groups the files of one
multi-select into one prompt. In chat, the agent’s ask_user tool asks the
question instead (naming the files in report_date_for); the reply is parsed
and applied on resume — the same rule as the endpoint below.
Moves every reading of that file to the date (date_source: manual) and
answers {"file_key", "report_date", "moved", "skipped"} — skipped counts
readings whose indicator already had a row on that date, which stay put rather
than overwrite it. Omit report_date to keep the upload time and stop the
prompt (date_confirmed: true). A file uploaded into someone else’s record
needs that member’s care-circle write grant; otherwise the file is reported
as not found.
Data Storage
Extracted indicator data is stored in the observation model (th_observation
and its coding tables, mirobody/schema/30_observations.sql). The write
goes through collect/observations.py:ingest — the one writer of those tables —
which freezes the extraction as read (th_extraction), stores every field as
printed, codes each row and skips a row the same file already wrote. A
deleted file’s rows are erased with it (the privacy path), so a re-upload
after a delete writes them fresh:
File Deletion
Endpoint
Request Body
Response Example
Cascade Deletion
When deleting files, the system automatically performs cascade deletion:- Storage Deletion: Delete file from object storage (S3/OSS)
- Database Update: Update file list in
th_messagestable - Health Data Cleanup: Erase the observations extracted from the file (
observations.erase, cascading to their coding and day authority) - Genetic Data Cleanup: If genetic file, delete its
th_genotype_setand cascadingth_genotyperows, plus any unmigratedth_series_data_geneticrows - Message Marking: If all files are deleted, mark message as deleted
API Reference
File Service Endpoints
Authentication
All endpoints require a valid authentication token:- REST API: Use
Authorization: Bearer <token>header - WebSocket: Pass via URL parameter
?token=<token>
Data Models
FileUploadData
FileDeleteRequest
FileProcessingResult
Error Handling
WebSocket Error Messages
REST API Error Response
Common Errors
Timeout Handling
WebSocket connections receive a notification on timeout:Best Practices
- Large File Uploads
- Use WebSocket upload with chunked transfer
- Recommended chunk size: 1MB
- Implement resumable upload mechanism
- Batch Uploads
- Keep uploads to a handful of files and moderate total size. (These are recommendations for client behavior — the server does not currently enforce a per-upload file count or total-size ceiling, so a client that ignores them fails slowly rather than being rejected.)
- Progress Monitoring
- Listen for
upload_progressmessages during WebSocket uploads - Handle
file_progressto display individual file progress
- Error Handling
- Implement retry mechanism (recommended: max 3 retries)
- Capture and display user-friendly error messages
- Connection Keep-alive
- Send ping every 30 seconds for WebSocket connections
- Handle pong response to confirm connection status