Private beta — onboarding design partners

Turning decades of materials science literature into ML-ready data

MatDB builds large-scale, multimodal datasets from published research — text, tables, images, and spectra — powering the next generation of materials design, discovery, and process optimization tools.

12,400+
Papers processed
3.8M
Data points extracted
14
Material classes covered
6
Active pilot partners
WORKING WITH
Partner OnePartner TwoPartner ThreePartner Four
"...the alloy exhibited a yield strength of 412 MPa at room temperature, with grain size refined to 2.3 μm after ECAP processing..."
extracted
PropertyValueUnit
Yield strength412MPa
Test condition298K
Grain size2.3μm
Technology

From a paragraph to a structured, ML-ready data point

Our extraction pipeline turns unstructured literature into validated, multimodal materials data — at a scale no manual process can match.

01

Ingest

Papers, patents, and theses pulled from journals, repositories, and partner archives.

02

Extract

Domain-specific entity recognition, table parsing, and figure and spectra digitization.

03

Fuse

Text, tabular, and image or spectral data aligned to the same material and experiment.

04

Validate

Human-in-the-loop QA against ground truth, with provenance tracked to source.

Domain-specific extraction

Materials nomenclature is inconsistent across papers, decades, and sub-fields — generic NER models miss most of it.

True multimodal fusion

We align text, tables, images, and spectra to the same underlying experiment, not just OCR on a page.

Validated, not scraped

Every extracted value is checked against reported ground truth before it enters the dataset.

Full provenance

Every data point traces back to its exact source publication, page, and figure or table.

Dataset

The dataset behind the models

Material classesAlloys, polymers, ceramics, battery materials, catalysts
ModalitiesText · tables · micrographs · spectra · XRD patterns
Source coverage140+ journals, 1985–present
Last updatedAugust 2026

Every entry is traced to its source publication and validated against reported ground truth by domain-trained reviewers before release.

Download a sample dataset

Get a representative slice of our extracted, structured data — no commitment required.

We respect publisher terms and maintain full attribution and provenance for every extracted data point. Read our data ethics policy →
Traction

Early validation from industry and research leaders

We're in active pilots with organizations solving real materials discovery and process optimization problems.

Industry pilot

Fortune 500 battery manufacturer

Using extracted electrolyte and cathode data to accelerate formulation screening ahead of physical testing.

Research partner

Leading academic materials lab

Building a validated alloy property dataset to train property-prediction models for high-throughput screening.

Private beta
Now
Design partner rollout
Q1 2027
General availability
Q3 2027
Team

Built by people who've lived this problem

Founder name

CO-FOUNDER, CEO

Ph.D. Materials Science. Previously led high-throughput screening at [lab]. Spent years re-typing values out of PDFs by hand.

Founder name

CO-FOUNDER, CTO

Ph.D. Machine Learning. Built information-extraction systems for scientific text prior to founding MatDB.

Founder name

CO-FOUNDER, HEAD OF DATA

Former research scientist focused on alloy design; led data infrastructure for a national lab consortium.

ADVISORS
Advisor name, TitleAdvisor name, TitleAdvisor name, Title
Company

Materials discovery is too slow. We think data is the bottleneck.

The problem

New materials typically take over a decade to go from discovery to deployment. Most of the relevant knowledge already exists — buried, unstructured, and unsearchable across millions of published papers.

Our bet

Structured, multimodal data unlocks AI tools that can meaningfully compress that timeline — but only if the underlying data is accurate, validated, and traceable to its source.

Where we're headed

From a data provider today to the default data and modeling layer materials teams build on — spanning discovery, formulation, and process optimization.

Contact

Let's talk

Investors

Reach out directly for our deck, data room access, or to schedule time with the founders.

matintelairoorkee@gmail.com

Partners & pilots

Interested in piloting MatDB's data or tools at your organization? We'd like to hear about your use case.

matintelairoorkee@gmail.com