Skip to main content
Jump by task

Before you write a provider: do you need one?

A provider integration is two very different things glued together — the OAuth dance, token storage, rate limits and pull windows (IO, per deployment) and the translation of a vendor’s JSON into standardised facts (a pure function of the payload). This project owns the second half properly and the first half only as far as its own reference application needs. If you need broad device access — eight or more vendors, a phone SDK, webhooks, multi-account fan-out — put open-wearables in front: twelve provider strategies, iOS/Android/Flutter/React Native SDKs, a developer portal, Celery sync. Then decode its API here (mirobody.kernel.decoders.open_wearables) and you get UCUM units, LOINC where the code is public, day windows that survive daylight saving, quality gates, corrections, medications and the two query tools on top. Run open-wearables to connect devices; run mirobody to make the data mean one thing and answer questions about it. Write a provider here when you need a vendor neither covers, or when the deployment must hold its own credentials.

What a provider is made of

A provider does not have to live in this repository. Ship it as a package that declares a mirobody.providers entry point pointing at the module with your BasePullProvider subclass, pip install it next to mirobody, and the pull loop picks it up at boot with no config change — or drop the mirobody_<slug>/ directory into a folder listed in PROVIDER_DIRS. Both paths call the same create_provider(config). examples/mirobody_example_plugin/ is a complete, installable example of the first.

  1. The decode table

A dict from the vendor’s field path to (catalogue metric, a function that converts INTO the catalogue's unit):
The unit is never written in the table. It comes from the catalogue, so a decoder cannot disagree with the aggregator about what ms means. The table says only how to get there. (The one bug this rule would have prevented in this repository: a kJ→kcal conversion applied the wrong way round, invisible because the number stayed plausible.) Three more rules:
  • Never fabricate a time. A record without one decodes to nothing. A synthetic timestamp is indistinguishable from a measured one afterwards.
  • Never guess a metric. A key the catalogue does not know returns nothing and is quarantined. metrics.mapping_for returning None is an answer.
  • Emit intervals as intervals. A night’s stages are spans, not one total; two syncs of the same night overlap and adding their totals counts it twice.
decode(data_type, item, tz, *, pulled_at_ms, source_record_id, ingested_at_ms) -> list[series.Fact], pure — no clock, no database, no config.

  1. The IO shell

format_data(self, fmt_input: FormatDataInput) -> StandardPulseData calls decoders.decode and wraps the facts with records_from_facts. Everything else in the class is authentication, pagination and storage. The credential state machine is mirobody.kernel.connect: the base pull loop already backs off after an authorization failure and stops after a threshold, so a changed password does not become an account lockout.

  1. The samples

mirobody/kernel/decoders/samples/<vendor>/*.json: payloads shaped like the vendor’s PUBLIC documentation, each with its expected facts computed by hand. Never by running the decoder — a sample generated from the code under test asserts that the code does what it does. They ship in the wheel, so a downstream repository can run mirobody.testing.samples.run_samples against its own decoders.

  1. Coverage — what the connector ACTUALLY carries

“This platform supports Oura” and “this platform brings you your blood pressure” are different claims, and a person choosing a device wants the second. connect.Coverage is generated from the decode table, so a page built from it cannot promise data the code does not produce. Regenerate with mirobody.testing.coverage.gen_coverage(decoders.coverage_matrix()) and embed with coverage.embed; coverage.stale is the CI check that fails the build when a decoder gains a metric and this table has not caught up. A metric listed here is one the decoder can EMIT. Whether a given person’s device records it is a further question, and no table can answer it.

The long-form walkthrough

Everything below is the full end-to-end guide: configuration, the OAuth flows, the core method reference, scopes, testing and troubleshooting. Start here if you are writing the IO shell rather than the decode table. This guide provides comprehensive instructions for integrating new device/service providers into the Mirobody Health platform.

Table of Contents

  1. Prerequisites
  2. Provider Architecture Overview
  3. Implementation Requirements
  4. Core Methods Reference
  5. Data Flow & Scopes
  6. Testing Your Provider
  7. Best Practices

  1. Prerequisites

Before integrating a new provider, ensure you have:

Technical Requirements

  • Python 3.12+ environment (pyproject.toml: requires-python = ">=3.12")
  • Access to the target device/service API documentation
  • OAuth credentials (OAuth1 or OAuth2) from the device vendor
  • Understanding of async/await patterns in Python
  • Familiarity with REST APIs and JSON data formats

Configuration Requirements

You’ll need to configure the following in your config.yaml:

Database Schema

Ensure your database has the provider-specific table:

  1. Provider Architecture Overview

Directory Structure

Class Hierarchy

Key Components

  1. OAuth Flow Handler: Manages user authentication
  2. Data Puller: Fetches data from vendor API
  3. Data Formatter: Transforms vendor data to standard format
  4. Database Service: Persists raw and formatted data

  1. Implementation Requirements

Step 1: Create Provider Directory

Two locations work, and the choice is about ownership rather than mechanism: a custom provider goes in a root providers/ directory (nothing in this repo to fork), a core one in mirobody/collect/providers/. The loader globs mirobody_*/provider_*.py in both.
The module file name need not match the directory slug — the shipped Garmin provider is mirobody_garmin_connect/provider_garmin.py. A mirobody_*/ directory with no provider_*.py fails test_installed.py rather than loading nothing silently. See mirobody/collect/providers/README.md.

Step 2: Define Provider Class

Your provider must inherit from BasePullProvider and implement all required methods:

Step 3: Define Data Mapping

Create a mapping between vendor API fields and standard indicators:

  1. Core Methods Reference

4.1 __init__(self)

Purpose: Initialize provider configuration and credentials Scope: Instance initialization Implementation:
Key Points:
  • Always call super().__init__() first
  • Use safe_read_cfg() for configuration values
  • Provide sensible defaults where possible
  • Validate critical configuration

4.2 create_provider(cls, config: Dict[str, Any]) -> Optional['YourProvider']

Purpose: Factory method for conditional provider instantiation Scope: Class method (called before instance creation) Implementation:
Key Points:
  • Returns None if provider cannot be initialized
  • Graceful failure - don’t raise exceptions
  • Log informational messages for debugging

4.3 info(self) -> ProviderInfo

Purpose: Provide metadata about the provider Scope: Property, accessed frequently for display/routing Implementation:
Key Points:
  • slug must be unique across all providers
  • auth_type must match your OAuth implementation
  • Logo should be hosted on CDN

4.4 link(self, request: Any) -> Dict[str, Any]

Purpose: Initiate OAuth authentication flow Scope: User-triggered, begins the linking process Implementation for OAuth2:
Implementation for OAuth1:
Key Points:
  • Store temporary OAuth state/tokens in Postgres temporary state with TTL
  • Include user_id mapping for callback retrieval
  • Return URL that user must visit
  • Handle both OAuth1 and OAuth2 flows appropriately

4.5 callback(self, *args, **kwargs) -> Dict[str, Any]

Purpose: Handle OAuth callback and complete authentication Scope: Triggered by OAuth provider redirect Implementation for OAuth2:
Implementation for OAuth1:
Key Points:
  • Retrieve stored state/tokens from Postgres temporary state
  • Exchange temporary tokens for permanent ones
  • Save credentials using appropriate db_service method
  • Trigger immediate data pull after successful linking
  • Clean up temporary Postgres temporary state keys

4.6 unlink(self, user_id: str) -> Dict[str, Any]

Purpose: Remove user connection and revoke access Scope: User-triggered, removes all provider data Implementation:
Key Points:
  • Attempt to revoke access at vendor (optional)
  • Always clean up local database regardless of vendor response
  • Log warnings for partial failures
  • Return success if database cleanup succeeds

4.7 format_data(self, fmt_input: FormatDataInput) -> StandardPulseData

Purpose: Transform vendor-specific data to standardized format. A pure transformation: everything that needs a database (the internal user id behind a vendor id, the user’s timezone) is resolved by the platform beforehand and arrives in fmt_input.context; the vendor payload is fmt_input.payload, untouched. That is what makes this method snapshot-testable against a recorded payload with no database at all. Scope: Called for every data batch received Implementation:
Key Points:
  • Always retrieve user timezone for accurate timestamps
  • Use data mapping configuration for consistency
  • Handle nested field navigation gracefully
  • Track processing metrics in processing_info
  • Return empty response on fatal errors

4.8 pull_from_vendor_api(self, *args, **kwargs) -> List[Dict[str, Any]]

Purpose: Fetch data from vendor API Scope: Called periodically or on-demand for data sync Implementation:
Key Points:
  • Implement pagination support
  • Handle rate limiting with retries
  • Use exponential backoff for transient errors
  • Support date range filtering
  • Return structured data for each data type

4.9 save_raw_data_to_db(self, raw_data: Dict[str, Any]) -> List[Dict[str, Any]]

Purpose: Persist raw vendor data to database Scope: Called before data formatting for audit trail Implementation:
Key Points:
  • Generate unique msg_id for deduplication
  • Use ON CONFLICT DO NOTHING to handle duplicates
  • Store complete raw payload as JSONB
  • Include both theta_user_id and external_user_id
  • Return augmented data with msg_id

4.10 is_data_already_processed(self, raw_data: Dict[str, Any]) -> bool

Purpose: Check if data has already been processed Scope: Called before saving/formatting to avoid duplicates Implementation:
Key Points:
  • Simple return False if using database constraints
  • Implement explicit checks only if needed
  • Log errors but don’t fail the pipeline

4.11 _pull_and_push_for_user(self, credentials: Dict[str, Any]) -> bool

Purpose: Unified pull and push workflow for a single user Scope: Called after OAuth linking or by scheduled tasks Implementation:
Key Points:
  • Handle token refresh for OAuth2
  • Pull recent data (e.g., last 2 days)
  • Push data through the pipeline asynchronously
  • Track success/error counts
  • Clean up invalid credentials on auth failure

  1. Data Flow & Scopes

Overall Data Flow

Scope Definitions

Public Methods (Called by Framework)

  • create_provider(): Class method, called during provider registration
  • info: Property, accessed for routing and display
  • link(): Endpoint-triggered by user action
  • callback(): Endpoint-triggered by OAuth redirect
  • unlink(): Endpoint-triggered by user action
  • format_data(): Pipeline-triggered for data transformation
  • save_raw_data_to_db(): Pipeline-triggered before formatting
  • is_data_already_processed(): Pipeline-triggered for deduplication

Private/Internal Methods

  • _pull_and_push_for_user(): Internal, triggered after linking or by scheduler
  • pull_from_vendor_api(): Internal, called by _pull_and_push_for_user
  • _process_*_data(): Internal helpers for format_data
  • _fetch_paginated_data(): Internal helper for API calls
  • get_valid_access_token(): Internal, OAuth2 token management
  • _generate_authorization_url(): Internal, OAuth flow helper
  • _handle_oauth_callback(): Internal, OAuth flow helper

Threading & Concurrency

  • All methods are async and use await for I/O operations
  • Use asyncio.create_task() to trigger background jobs
  • Use asyncio.gather() for concurrent API calls
  • Use asyncio.Semaphore() to limit concurrent requests

  1. Testing Your Provider

Unit Testing

Create test_provider_<provider>.py:

Integration Testing

Manual Testing

  1. OAuth Flow:
  2. Data Pull:
  3. Data Formatting:

Testing Checklist

  • Provider can be instantiated with valid config
  • Provider returns correct metadata via info
  • link() generates valid OAuth URL
  • callback() successfully exchanges tokens
  • Credentials are saved to database correctly
  • unlink() removes credentials from database
  • pull_from_vendor_api() fetches data successfully
  • save_raw_data_to_db() persists raw data
  • format_data() transforms data correctly
  • All StandardIndicators are mapped correctly
  • Timezones are handled properly
  • Rate limiting is handled gracefully
  • Token refresh works (OAuth2)
  • Error handling is robust
  • Logging is comprehensive

  1. Best Practices

Configuration Management

  1. Use Environment-Specific Configs:
  2. Validate Configuration on Startup:
  3. Provide Sensible Defaults:

Error Handling

  1. Log All Errors with Context:
  2. Fail Gracefully:
  3. Retry Transient Errors:

Data Mapping

  1. Use Configuration-Driven Mapping:
  2. Handle Missing Fields:
  3. Validate Data Types:

Performance Optimization

  1. Use Concurrent Requests:
  2. Implement Pagination:
  3. Cache Expensive Operations:

Security

  1. Never Log Sensitive Data:
  2. Validate Input:
  3. Use Secure Token Storage:

Monitoring & Observability

  1. Track Key Metrics:
  2. Use Structured Logging:
  3. Implement Health Checks:

Appendix A: StandardIndicator Reference

Common indicators you’ll map to:

Activity Indicators

  • STEPS: Step count
  • DISTANCE: Distance traveled (meters)
  • CALORIES_ACTIVE: Active calories burned
  • CALORIES_BASAL: Basal metabolic rate calories

Heart Rate Indicators

  • HEART_RATE: Heart rate (bpm)
  • HEART_RATE_MAX: Maximum heart rate
  • RESTING_HEART_RATE: Resting heart rate
  • HRV_RMSSD: Heart rate variability (RMSSD in ms)

Sleep Indicators

  • SLEEP_IN_BED: Time in bed (milliseconds)
  • SLEEP_ANALYSIS_AWAKE: Time awake during sleep (ms)
  • SLEEP_ANALYSIS_ASLEEP_CORE: Light sleep time (ms)
  • SLEEP_ANALYSIS_ASLEEP_DEEP: Deep sleep time (ms)
  • SLEEP_ANALYSIS_ASLEEP_REM: REM sleep time (ms)
  • SLEEP_EFFICIENCY: Sleep efficiency (percentage)
  • SLEEP_DISTURBANCES: Number of sleep disturbances

Body Metrics

  • WEIGHT: Body weight (kg)
  • HEIGHT: Height (meters)
  • BMI: Body mass index
  • BODY_FAT_PERCENTAGE: Body fat percentage
  • BLOOD_OXYGEN: SpO2 percentage

Workout Indicators

  • WORKOUT_DURATION_LOW: Low intensity duration (minutes)
  • WORKOUT_DURATION_MEDIUM: Medium intensity duration (minutes)
  • WORKOUT_DURATION_HIGH: High intensity duration (minutes)
  • ALTITUDE_GAIN: Altitude gained (meters)
  • SPEED: Average speed (m/s)

Appendix B: Common Issues & Solutions

Issue: OAuth Callback Not Working

Symptoms: Callback returns 404 or fails silently Solutions:
  1. Verify redirect URL matches exactly in vendor dashboard
  2. Check Postgres connectivity: OAuth state lives in th_ephemeral
  3. Ensure the provider is loaded, so its callback route is registered: its directory must be on PROVIDER_DIRS (config.yaml)
  4. Verify state parameter is URL-encoded properly

Issue: Token Expired During Data Pull

Symptoms: 401 errors from vendor API Solutions:
  1. Implement token refresh (OAuth2)
  2. Check expires_at calculation
  3. Verify refresh token is saved correctly
  4. Test get_valid_access_token() method

Issue: Data Not Formatting Correctly

Symptoms: Empty health records or missing indicators Solutions:
  1. Check data mapping configuration
  2. Verify field names match vendor API response
  3. Log raw data to inspect structure
  4. Test converter functions separately
  5. Ensure timezone is retrieved correctly

Issue: Rate Limiting

Symptoms: 429 errors from vendor API Solutions:
  1. Implement exponential backoff
  2. Respect Retry-After header
  3. Reduce concurrent request limit
  4. Cache frequently accessed data

Issue: Duplicate Data

Symptoms: Same data processed multiple times Solutions:
  1. Ensure msg_id is unique and consistent
  2. Add database constraint: UNIQUE(msg_id)
  3. Implement is_data_already_processed() check
  4. Use ON CONFLICT DO NOTHING in insert queries

Support & Resources

  • Example providers (read these first — they are the real thing, not samples):
    • Garmin: mirobody/collect/providers/mirobody_garmin_connect/provider_garmin.py
    • Whoop: mirobody/collect/providers/mirobody_whoop/provider_whoop.py
    • Oura: mirobody/collect/providers/mirobody_oura/provider_oura.py
  • Platform internals: mirobody/collect/providers/_platform/ — base.py is the contract you implement, platform.py does discovery and pull scheduling.
  • Testing: docs/testing.md. The maintainers’ internal suite (not published in this repository) additionally snapshots format_data() output for every shipped provider.
  • Data contract: mirobody/collect/ingest/models/requests.py — StandardPulseData and friends, the shape every provider must produce.
For questions or assistance, contact the platform team or create an issue in the repository.
Maintenance note: this guide is prose without a test gate — when it and the code disagree, the code wins. The shipped providers listed above are the executable reference.