Skip to main content

Overview

This module provides comprehensive health data file processing capabilities, including file upload, health indicator extraction, and file deletion. The system uses a modular design, supports multiple file types, and provides real-time progress feedback.

Table of Contents


Features

  • ✅ Multiple Upload Methods: Supports WebSocket real-time upload and REST API batch upload
  • ✅ Real-time Progress Feedback: WebSocket connections provide real-time progress updates for file upload and processing
  • ✅ Smart File Recognition: Automatically identifies file types and selects the appropriate handler
  • ✅ Health Indicator Extraction: Uses LLM to automatically extract health indicator data from medical reports
  • ✅ Multi-format Support: Supports PDF, images, Office documents, text, genetic data and more
  • ✅ PDF Parallel Processing: Multi-page PDFs are processed in parallel for improved efficiency
  • ✅ File Summary Generation: Automatically generates file content summaries
  • ✅ Cascade Deletion: Automatically cleans up associated health data when files are deleted

Supported File Types

Two lists, and they are not the same list. mirobody/utils/file_types.py is the UPLOAD accept-list: what a user is allowed to hand the server. It carries .xls and .xlsb, which no parser here reads. What a file can be READ as is mirobody/documents/detect.py: EXTRACTABLE_SUFFIXES (15, through a parser or OCR) plus TEXT_EXTENSIONS (8, decoded directly) is the 23 the README counts. When this table and that module disagree, the module is the contract. Whatever the handler, the TEXT of a document comes from one place: mirobody/documents/ (detect.kind by extension, content type and, when those lie, the bytes; extract.extract_text by kind). A PDF gives up its embedded text layer page by page and only the pages that have none — scans — are rendered and handed to the vision model, one image at a time; a photo is downscaled and OCR’d; a spreadsheet or Word file never reaches a model at all. The result is cached by content hash through th_files, so the same bytes are never OCR’d twice, and the file summary is generated from that text rather than from the file.

File Upload

WebSocket upload provides real-time progress feedback, ideal for large file uploads and scenarios requiring real-time status updates.

Endpoint

Connection Flow

upload_completed and extraction_completed answer different questions, and a client needs both: the first says the file arrived and is readable, the second says whether anything could be extracted from it. Extraction runs after the upload has already completed, so a client that stops listening at upload_completed will render a stored file as a finished one.

Message Types

Client sends:
Server response:
Client sends:
Server response:
Sent as soon as the report’s own date has been probed — seconds after upload, while the 15-25 s indicator extraction is still running — so the client can ask about a missing date without waiting for the whole extraction.
date_source is document when the date was read off the report itself, and upload when none was found and the upload time is standing in as a placeholder. A client should offer to correct the second case. This is the event that says whether extraction worked, and the only one that does. upload_completed above reports that the FILE arrived and is stored — it says 1 files successful even when every model call failed, because the file is genuinely stored and readable either way.
failed: true means extraction could not run — no provider is configured, or every call to the configured one errored. The sentence explaining which is on the file row (error, from GET /api/v1/health-indicators/files); the file itself is stored and readable. failed: false with indicators_count: 0 is a different and ordinary outcome: a document that genuinely has no indicators in it. A client must not render the first two like the third — that is what made a green row over an empty list, with the cause visible only in server logs. Client sends:
Server response:

Timeout Mechanism


REST API Upload

REST API provides a simple file upload method, suitable for simple scenarios or applications that don’t require real-time progress.

Endpoint

Request Parameters

Response Example

Response Codes


Processing Flow

Architecture Overview

Processing Steps

  1. File Upload Phase (0-30%)

  • Receive file data
  • Validate file type and size
  • Generate unique file identifier
  • Upload to object storage (S3/OSS)

  1. File Type Recognition (30-35%)

The system automatically identifies file types via FileHandlerFactory:

  1. Content Processing Phase (35-90%)

PDF File Processing:
Image File Processing:

  1. Summary Generation Phase (90-95%)

  • Generate file content summary using LLM
  • Generate intelligent file name

  1. Result Saving Phase (95-100%)

  • Save processing results to database
  • Write the extracted indicators as observations (th_observation, coded on the way in)
  • Update user health profile

Health Indicator Extraction

Description

The system uses Large Language Models (LLM) to automatically identify and extract health indicator data from medical reports.

Configuration

Supported Indicator Types

  • Complete Blood Count (WBC, RBC, Hemoglobin, etc.)
  • Biochemistry (Liver function, Kidney function, Lipids, etc.)
  • Physical Examination (Blood pressure, Heart rate, Weight, etc.)
  • Tumor Markers
  • Thyroid Function
  • Other Medical Test Indicators

Extraction Result Format

Report Date

content_info.date_time becomes the start_time/end_time of every reading extracted from the file. When the document shows no date (or one the parser cannot read), the readings are filed under the user’s current time — and that fallback is labelled, not silent: each reading’s comment JSON and the file row carry date_source: The labelling exists because extraction runs one file at a time. A report photographed as three screenshots shows its date on the first page only, and “page 2 of the same report” is indistinguishable from “a second report whose date did not come out” — so nothing is inherited automatically. The files listing (GET /api/v1/data/uploaded-files) exposes report_date, date_source and date_confirmed per file; the web client asks about upload_time files and offers the dates read from the other files of the same upload (same created_source_id, or uploaded within a few minutes of it — the web client opens one upload session per file).
The date is looked up before the indicators, in its own small model call, and the upload WebSocket carries two events after upload_completed: report_date_detected (file_key, report_date, date_source, seconds after the upload) and extraction_completed (adds indicators_count); both carry the upload’s sessionId, which is how the web client groups the files of one multi-select into one prompt. In chat, the agent’s ask_user tool asks the question instead (naming the files in report_date_for); the reply is parsed and applied on resume — the same rule as the endpoint below. Moves every reading of that file to the date (date_source: manual) and answers {"file_key", "report_date", "moved", "skipped"} — skipped counts readings whose indicator already had a row on that date, which stay put rather than overwrite it. Omit report_date to keep the upload time and stop the prompt (date_confirmed: true). A file uploaded into someone else’s record needs that member’s care-circle write grant; otherwise the file is reported as not found.

Data Storage

Extracted indicator data is stored in the observation model (th_observation and its coding tables, mirobody/schema/30_observations.sql). The write goes through collect/observations.py:ingest — the one writer of those tables — which freezes the extraction as read (th_extraction), stores every field as printed, codes each row and skips a row the same file already wrote. A deleted file’s rows are erased with it (the privacy path), so a re-upload after a delete writes them fresh:

File Deletion

Endpoint

Request Body

Response Example

Cascade Deletion

When deleting files, the system automatically performs cascade deletion:
  1. Storage Deletion: Delete file from object storage (S3/OSS)
  2. Database Update: Update file list in th_messages table
  3. Health Data Cleanup: Erase the observations extracted from the file (observations.erase, cascading to their coding and day authority)
  4. Genetic Data Cleanup: If genetic file, delete its th_genotype_set and cascading th_genotype rows, plus any unmigrated th_series_data_genetic rows
  5. Message Marking: If all files are deleted, mark message as deleted

API Reference

File Service Endpoints

Authentication

All endpoints require a valid authentication token:
  • REST API: Use Authorization: Bearer <token> header
  • WebSocket: Pass via URL parameter ?token=<token>

Data Models

FileUploadData

FileDeleteRequest

FileProcessingResult


Error Handling

WebSocket Error Messages

REST API Error Response

Common Errors

Timeout Handling

WebSocket connections receive a notification on timeout:

Best Practices

  1. Large File Uploads

  • Use WebSocket upload with chunked transfer
  • Recommended chunk size: 1MB
  • Implement resumable upload mechanism

  1. Batch Uploads

  • Keep uploads to a handful of files and moderate total size. (These are recommendations for client behavior — the server does not currently enforce a per-upload file count or total-size ceiling, so a client that ignores them fails slowly rather than being rejected.)

  1. Progress Monitoring

  • Listen for upload_progress messages during WebSocket uploads
  • Handle file_progress to display individual file progress

  1. Error Handling

  • Implement retry mechanism (recommended: max 3 retries)
  • Capture and display user-friendly error messages

  1. Connection Keep-alive

  • Send ping every 30 seconds for WebSocket connections
  • Handle pong response to confirm connection status

Code Examples

JavaScript WebSocket Upload

Python REST API Upload