Extract multiple FHIR resources from a document
/lang2fhir/document/multiExtracts text from a PDF, image, RTF, or XML/C-CDA document and converts it into multiple FHIR resources, returned as a transaction Bundle. Combines document text extraction with multi-resource detection. Automatically detects Patient, Condition, MedicationRequest, Observation, and other resource types. Resources are linked with proper references (e.g., Conditions reference the Patient).
Patient identifier handling. US Core requires Patient.identifier (a business identifier such as an MRN). When the source text contains an identifier, it is extracted with an appropriate URI system. When the source text does not contain a detectable identifier, a synthetic one is generated with system: "urn:phenoml:lang2fhir-generated-id" and a UUID value so the bundle remains FHIR-valid and US Core conformant. Callers who need a tenant-specific namespace should rewrite the synthetic system after extraction.
Split classifications (optional). config.split_classifications is a caller-defined list, not a fixed taxonomy. Choose each classification id and write a natural-language description for the per-page classifier. For each page, the classifier assigns the best-matching classification or leaves the page ungrouped. Classifications with operation: "group" keep matching pages and label resources extracted from those pages; classifications with operation: "drop" remove matching pages before extraction. The clinical and admin ids in the example are illustrative, not a fixed set.
Body parameters
FHIR version to use
Base64 encoded file content.
Supported file types: PDF (application/pdf), PNG (image/png), JPEG (image/jpeg), TIFF (image/tiff), RTF (application/rtf), XML/C-CDA (text/xml).
TIFF, RTF, and XML/C-CDA uploads are available on dedicated instances only.
File type is auto-detected from content magic bytes.
The decoded file must not exceed 20 MiB. RTF and XML/C-CDA documents whose extracted text exceeds 1 MiB are rejected.
Generic XML must include an XML declaration; C-CDA documents rooted at ClinicalDocument may omit it.
Optional FHIR provider name for provider-specific profiles
Partial context for the patient the document is primarily about. This is not a complete FHIR Patient resource. Lang2FHIR uses the available context to identify a generated primary Patient reliably. An identifier supplied here is added to that Patient; when no Patient is generated, it is used in logical references on generated clinical resources.
Business identifier for the document's primary patient. When Lang2FHIR identifies that Patient in generated output, it adds this identifier to the Patient's identifier list (preserving existing identifiers). If no Patient is generated, Lang2FHIR uses it in logical references on generated clinical resources. Supply the patient-level identifier (not an order or specimen identifier).
Identifier system (namespace) for the patient identifier.
Identifier value for the patient within that system.
The known portions of the primary patient's name. Provide a non-empty family name or at least one non-empty given name.
Family name.
Given names. Matching succeeds when a generated name has a supplied given name.
Complete date of birth in YYYY-MM-DD format.
Administrative gender. This corroborates another match but does not identify a patient alone.
malefemaleotherunknownBusiness identifier for the document's primary patient. When Lang2FHIR identifies that Patient in generated output, it adds this identifier to the Patient's identifier list (preserving existing identifiers). If no Patient is generated, Lang2FHIR uses it in logical references on generated clinical resources. Supply the patient-level identifier (not an order or specimen identifier).
Identifier system (namespace) for the patient identifier.
Identifier value for the patient within that system.
Custom Implementation Guide name. When specified, profiles from this IG are included alongside the default profiles during resource detection. Default profiles are always the base layer; custom IG profiles are additive.
Deprecated; use the default 'standard' value. This field will be removed in a future release. 'standard' runs detection once; 'deep' runs detection multiple times for higher recall.
standarddeepFHIR validation method to apply to the generated bundle. 'none' skips validation (default). 'check' runs the bundle through a FHIR structure validator and includes the results in the response. 'fix' runs validation and attempts to auto-correct errors using an LLM (up to 3 validation passes). The response includes results from each pass. Warning: 'fix' can significantly increase latency due to multiple LLM and validation round-trips.
nonecheckfixOptional processing configuration shared across document endpoints.
Deprecated. Use split_classifications instead.
Optional per-page split classifications. Mutually exclusive with page_filter. This is a caller-defined list, not a fixed taxonomy: choose each classification id and write a natural-language description for the per-page classifier. For each page, the classifier assigns the best-matching classification or leaves the page ungrouped. Pages matching operation=drop are removed before extraction. Pages matching operation=group are kept, and extracted resources attributed to those pages include the classification id in response metadata and FHIR meta.tag. Example ids such as clinical and admin are illustrative, not a fixed set.
Unique classification id. The reserved id "ungrouped" is not allowed.
Natural-language description of pages that belong to this classification.
Opt-in faithfulness audit (honored by /lang2fhir/create/multi and /lang2fhir/document/multi). For each selected resource type an LLM checks whether selected dates and clinical code concepts are actually supported by the full source document. An unsupported individual coding is removed when another coding remains in its concept. Resources with an unsupported structural field, a profile-required coding, or no coding remaining in an affected concept, are pulled out of the returned bundle and reported under resource_review in the response.
The resource types to audit and which date or clinical code field kinds to check.
Successfully extracted FHIR resources from document
Response fields