A medical training grade corpus, institution authored and rights clean. Each record originates from the institution of publication itself, with named investigators, embargo timing, and aligned multimedia preserved at the record level.
Records in this corpus are captured at the first point — the institutional release — where provenance, attribution, and timing remain intact. Downstream collections typically encounter the same findings only after that source signal has been lost.
The properties that distinguish biomedical research — named authorship, institutional review, embargo timing, and editorial discipline — are precisely the properties that web-derived collections tend to lose along the way. Once those signals are stripped, they are difficult to reconstruct.
A corpus cannot convey embargo, retraction, or peer-review status if those signals were not preserved when the material was first collected.
This corpus is organised around the opposite premise: a documented source of record — the named institution, the named investigator, the embargo date — preserved at the level of the individual record, throughout a continuous thirty-year archive.
Each row describes a property that is commonly absent when biomedical material is collected from the open web, the reason it tends to be lost, and how this corpus preserves it from the point of original release.
These are not features of scale. Each reflects a structural property of how the material was produced — at the source, under embargo, with attribution intact. A deduplicated attribution graph of 48,051 unique experts and 51,737 expert–institution relationships (across 397 institutions in the medical expert-graph) connects records to named authorities.
Records originate from the named institution of publication, with researcher names, titles, and affiliations preserved at the record level for verifiable attribution.
Researcher- and PIO-approved images and video paired to release text at the record level — clean image–text and video–text signal with no post-hoc captioning.
Thirty years of continuous, embargo-dated provenance enables time-aware training that separates established consensus from emerging findings without conflation.
Institution-authored releases precede downstream coverage, giving the corpus a measurable lead over the aggregated medical news that later follows it.
A license is most fairly measured against the alternative: originating thirty years of institutional relationships, editorial discipline, and rights documentation from a standing start. The relevant comparison is time and feasibility.
Four characteristics of the current environment make a rights-reviewed, institution-authored biomedical layer a substantive consideration for organisations building serious medical systems.
Provenance, attribution, and temporal signal are increasingly recognised as distinct from raw volume. A system grounded in this material can cite, date, and attribute its claims in a way that web-derived corpora do not readily support.
As training-data provenance moves toward becoming an audited surface under emerging AI governance, a documented institution-to-platform chain is increasingly relevant to compliance and review.
The 2020–2025 cohort — a particularly active period in biomedical communication — is finite and already captured. The window itself is no longer open to be re-originated.
The medical cohort can be licensed in defined slices, and the structure supports category-, industry-, or temporal-exclusive arrangements where that serves both sides of an engagement.
Medicine and the health sciences are the strongest, most actively maintained area of the archive — produced under embargo, written for accuracy, tied to named researchers at named institutions across oncology, cardiovascular disease, neuroscience, infectious disease, and public health.
Named institution and investigator preserved per record, for attribution-grounded generation.
78,520 cohort multimedia assets — including 5,223 videos — approved at source and paired to release text. (126,522 archive-wide.)
Embargo-dated release timing allowing time-aware reasoning without conflation of recent and historical findings.
First-public-mention windows recording the corpus's lead over downstream reporting.
A license is shaped to the requirements of the system being built — rights, scope, attestation, and pathway, specified in dialogue with technical and legal counsel rather than presented as a fixed package.
A license is defined as a combination across seven dimensions. The corpus is structured to support different specifications depending on the system being built. A medical-cohort license is one combination among many — the first commercial expression of the architecture.
Each dimension is licensed independently. Specific structures — and any exclusivity across content domains, intelligence layers, or time windows — are defined in dialogue with technical and legal counsel, in light of the system the licensee is building.
A diligence package prepared for technical and legal counsel to examine the corpus at the record level. Five components, delivered together, each addressing a question a reviewing team would reasonably want to answer.
The contractual basis under which institutional content reaches the corpus, and the structure of the chain of custody from institution to platform.
A representative roster of contributing institutions — named, verifiable, and substantiated against the volume figures presented.
Field-level schema for records, expert graph, institution graph, and multimedia pairings — what arrives, how it is shaped, and how it is keyed.
A curated sample drawn from the medical cohort — real records with real attribution, embargo dates, and paired multimedia, for direct technical inspection.
A working subset structured for the licensee's own benchmarks — provenance-aware fine-tuning, citation grounding, and multimodal alignment — under terms that credit the evaluation forward into production.