MatDB builds large-scale, multimodal datasets from published research — text, tables, images, and spectra — powering the next generation of materials design, discovery, and process optimization tools.
| Property | Value | Unit |
|---|---|---|
| Yield strength | 412 | MPa |
| Test condition | 298 | K |
| Grain size | 2.3 | μm |
Our extraction pipeline turns unstructured literature into validated, multimodal materials data — at a scale no manual process can match.
Papers, patents, and theses pulled from journals, repositories, and partner archives.
Domain-specific entity recognition, table parsing, and figure and spectra digitization.
Text, tabular, and image or spectral data aligned to the same material and experiment.
Human-in-the-loop QA against ground truth, with provenance tracked to source.
Materials nomenclature is inconsistent across papers, decades, and sub-fields — generic NER models miss most of it.
We align text, tables, images, and spectra to the same underlying experiment, not just OCR on a page.
Every extracted value is checked against reported ground truth before it enters the dataset.
Every data point traces back to its exact source publication, page, and figure or table.
| Material classes | Alloys, polymers, ceramics, battery materials, catalysts |
| Modalities | Text · tables · micrographs · spectra · XRD patterns |
| Source coverage | 140+ journals, 1985–present |
| Last updated | August 2026 |
Every entry is traced to its source publication and validated against reported ground truth by domain-trained reviewers before release.
Get a representative slice of our extracted, structured data — no commitment required.
We're in active pilots with organizations solving real materials discovery and process optimization problems.
Using extracted electrolyte and cathode data to accelerate formulation screening ahead of physical testing.
Building a validated alloy property dataset to train property-prediction models for high-throughput screening.
Ph.D. Materials Science. Previously led high-throughput screening at [lab]. Spent years re-typing values out of PDFs by hand.
Ph.D. Machine Learning. Built information-extraction systems for scientific text prior to founding MatDB.
Former research scientist focused on alloy design; led data infrastructure for a national lab consortium.
New materials typically take over a decade to go from discovery to deployment. Most of the relevant knowledge already exists — buried, unstructured, and unsearchable across millions of published papers.
Structured, multimodal data unlocks AI tools that can meaningfully compress that timeline — but only if the underlying data is accurate, validated, and traceable to its source.
From a data provider today to the default data and modeling layer materials teams build on — spanning discovery, formulation, and process optimization.
Reach out directly for our deck, data room access, or to schedule time with the founders.
Interested in piloting MatDB's data or tools at your organization? We'd like to hear about your use case.