Hancom Product User Guide

The Start of RAG/LLM Preprocessing: Hancom Data Loader Document Parsing Solution – Transforming Documents into AI Training Data

HANCOM

Hancom Data Loader is a document parsing solution developed by Hancom. It is a document data preprocessing technology specialized for RAG-based AI training that helps AI understand HWP, HWPX, PDF, and OOXML.

Before an agent can act, Hancom Data Loader converts documents into AI-readable data, separating and extracting content down to semantic units and providing it as agent-friendly metadata.

Why is document datafication so difficult when adopting enterprise AI?

Because large language models (LLMs) cannot read internal documents such as PDF, HWP, and HWPX as-is.

This unstructured document processing is the common barrier faced by public institutions, large-enterprise DX organizations, and SI partners considering in-house AI assistants, retrieval-augmented generation (RAG), and agent adoption.

Why Document Structuring Is Necessary: The Gap Between LLMs and Unstructured Documents

Enterprises hold vast amounts of materials—manuals, reports, contracts, and official documents in HWP/HWPX—but not many are in a form that LLMs can use directly.

In PDFs, table, multi-column, and caption structures disappear, and with HWP/HWPX, meaning breaks with simple text extraction. This is why table rows and columns get mixed up and why answers can produce plausible summaries with missing evidence. Only by converting documents into structured data can LLMs properly read internal information.

Why Document Parsing Quality Matters in a RAG Pipeline

For organizations driving AI transformation (AX), structured document data is an asset that affects not only retrieval but also the reliability of agent decisions and execution.

This is because document preprocessing quality in a RAG pipeline determines the entire process—retrieval, reasoning, and execution.

If table structures are lost during parsing, irrelevant text can be mixed in and mapped to the wrong locations during embedding.

Hancom Data Loader, Developed by Hancom: An Automated Document Structure Analysis Solution

Hancom Data Loader, a Document Structure Analysis Parsing Solution: Definition and Core Role

Hancom Data Loader: An Automated Document Structure Analysis Solution

Hancom Data Loader is a document parsing solution developed by Hancom. It converts various document formats such as HWP, HWPX, PDF, and OOXML into structured data that AI can understand, and can be used to build RAG-based knowledge search systems, secure training data for Large Language Models (LLMs), and digitize corporate documents.

It separates and extracts document semantic units—including tables, titles, captions, and hierarchical structures—based on OCR (Optical Character Recognition) extraction and provides them as agent-friendly metadata. As a document structure-preserving extraction method rather than general OCR, it is establishing itself as a core preprocessing technology for RAG construction.
Document analysis is performed based on Document Layout Analysis (DLA), OCR, and Table Structure Recognition (TSR). It distinguishes document components such as text, tables, and images, and extracts structural information together. Its key feature is that it goes beyond simple text extraction to separate semantic units within the document and provide them in metadata format.

In addition, HWP and HWPX documents provide paragraph-based hierarchical structure information, and PDF_AI supports visual information interpretation functions based on Image Captioning.

(※ The image captioning function is currently in the PoC stage, and the commercial release schedule will be announced later.)

From Document Structure Analysis to Data Extraction

Hancom Data Loader generates structured data through a three-step processing flow: document input → document analysis → data extraction. The extracted data can then be linked with RAG pipelines and LLMs to build AI services.

In the document analysis stage, it runs two engines in parallel: rules-based analysis and AI-based document structure understanding. Structured documents with consistent formats are handled with rules, while unstructured documents that mix figures, tables, and images are analyzed precisely with AI.

It provides, in a single processing flow, the core capabilities required for document preprocessing—recognizing diverse document elements, detecting multi-column reading order, and extracting complex table and chart structures.

The extracted structured data can be directly integrated with Hancom’s proprietary RAG solution, HancomPedia, enabling rapid deployment of a document-based AI search system.

Image explaining the AI document processing and data utilization process
A four-step pipeline combining Hancom Data Loader and HancomPedia, processing the entire workflow from document input to utilization within a single solution

Document Data Extraction Targets: Support for Multiple Document Formats

Core Format Support for Public-Sector AI Adoption: HWP/HWPX Parsing

As the government moves to expand HWPX usage from 2026 and gradually restrict HWP attachments, full support for HWP and HWPX is becoming a prerequisite for public-sector AI adoption. Hancom Data Loader supports all HWP versions from 3.0 onward, and with a parser developed in-house by Hancom, it preserves more than 20 layout elements without loss without converting to PDF.

Universal Format Support: PDF and OOXML (DOCX, XLSX, PPTX)

PDF(AI) handles layout analysis and table structure restoration through Document Layout Analysis (DLA), Table Structure Recognition (TSR), and OCR, and minimizes data loss for scanned documents through OCR integration. OOXML (DOCX, XLSX, PPTX) extracts text from Office documents using a text extraction method.

Unstructured Visual Documents: PNG/JPG

PNG and JPG automatically recognize text and layout to datafy even non-digitized materials such as fax documents, scanned images, and photographed documents.

👉Go to Hancom Data Loader Live Demo

Hancom Data Loader Key Features – Document Structure Analysis based on DLA, TSR, and OCR

Document Layout Analysis (DLA) – AI-based structural analysis that extracts even metadata

Hancom Data Loader’s Document Layout Analysis (DLA) classifies objects into categories and extracts tables, footnotes, images, charts, captions, formulas, and even fonts as metadata.

When metadata is extracted together, RAG retrieval accuracy improves, enabling policy search, administrative QA, and internal knowledge agents to use highly reliable data.

Table, Image, and Chart Structure Recognition – Table Structure Recognition (TSR) technology, the core of unstructured data processing

Key information in corporate documents is often concentrated in tables and charts rather than the main text. Based on Table Structure Recognition (TSR), Hancom Data Loader restores cell relationships—including borderless tables, merged cells, and tables within tables—to convert them into Markdown or HTML, and improves AI result quality by preserving multi-column layouts.

Data Separation and Post-Editing: Data Loader Studio

We provide Data Loader Studio (an extension solution) for documents where it is difficult to ensure accuracy through automatic extraction alone. Extraction results preprocessed with Data Loader can be reviewed by comparing them side-by-side with the original document. Users can directly check the hierarchy, category, and reading order, and post-edit necessary parts to ensure that the document structure and semantic relationships are refined to suit actual business purposes. In addition, through semantic-based tagging and labeling, extraction criteria can be supplemented to fit the customer’s document format and refined into data quality suitable for training and searching.

Image explaining the Hancom Data Loader document utilization flow

The Starting Point for Document-Based AI: Hancom Data Loader

🖥️ Hancom Data Loader

Hancom Data Loader is a document parsing solution that converts HWP, HWPX, PDF, and OOXML into structured data. The extracted data can be integrated with our RAG solution, HancomPedia, enabling a single Hancom stack from document collection to search and answers. Start with document preprocessing that agents can trust—start with Hancom Data Loader.

👉 Go to Hancom Data Loader Live Demo

👉Inquire about Hancom Data Loader Implementation


References

1) Yonhap News, “”Reducing hwp that AI can’t read”… Government to transition to open hwpx,” 2026