fig3

LLM-driven materials knowledge extraction: multimodal parsing, ontology, and agentic systems

Figure 3. Two document-parsing paradigms for materials-science papers. (A) Modular OCR- and rule-based pipelines separate layout, table, chart, and chemical-structure extraction; (B) Unified end-to-end multimodal LLM parsing directly maps page images or parsed blocks to machine-readable structured output. The two routes differ in controllability, error propagation, cost, and traceability. OCR: Optical character recognition; LLM: large language model.

Journal of Materials Informatics
ISSN 2770-372X (Online)
Follow Us

Portico

All published articles are preserved here permanently:

https://www.portico.org/publishers/oae/

Portico

All published articles are preserved here permanently:

https://www.portico.org/publishers/oae/