The Start of RAG/LLM Preprocessing: Hancom Data Loader Document Parsing Solution – Transforming Documents into AI Training Data
Hancom Data Loader is a document parsing solution developed by Hancom. It is a document data preprocessing technology specialized for RAG-based AI training that helps AI understand HWP, HWPX, PDF, and OOXML.
Before an agent runs, Hancom Data Loader converts documents into data that AI can understand, separating and extracting content down to semantic units and providing it as agent-friendly metadata.
Why is document datafication so difficult when adopting enterprise AI?
This is because large language models (LLMs) cannot read internal documents such as PDF, HWP, and HWPX as-is.
A common challenge faced by public institutions, enterprise DX teams, and SI partners considering the adoption of in-house AI assistants, RAG (Retrieval Augmented Generation), or agents is handling unstructured documents.
Why Document Structuring Is Necessary: The Gap Between LLMs and Unstructured Documents
Enterprises hold vast amounts of materials—manuals, reports, contracts, and official documents in HWP/HWPX—but not many are in a form that LLMs can use directly.
In PDFs, table, multi-column, and caption structures can be damaged, and in HWP and HWPX files, extracting only text does not convey meaning accurately. As table row and column order changes, answers can differ from the actual content. Only by converting documents into structured data can LLMs properly read internal information.
Why Document Parsing Quality Matters in a RAG Pipeline
For organizations pursuing AI transformation (AX), structured document data affects not only search but also the reliability of agent judgment and execution.
In a RAG pipeline, document preprocessing quality is directly reflected in retrieval, decision-making, and execution results.
If table structure information is lost during parsing, irrelevant text may be included, and during embedding, positions can be distorted and mapped incorrectly.

Document parsing solution: definition and core role of Hancom Data Loader
Hancom Data Loader: an automated document parsing solution
Hancom Data Loader is a document parsing solution developed by Hancom. It converts various document formats such as HWP, HWPX, PDF, and OOXML into structured data that AI can understand, and can be used to build RAG-based knowledge search systems, secure training data for Large Language Models (LLMs), and digitize corporate documents.
It separates and extracts document semantic units—including tables, titles, captions, and hierarchical structures—along with OCR (Optical Character Recognition)–based extraction, and provides them as agent-friendly metadata. As a document structure-preserving extraction method rather than general OCR, it is establishing itself as a core preprocessing technology for building RAG systems.
It analyzes documents based on DLA (Document Layout Analysis), OCR, and TSR (Table Structure Recognition), distinguishing document components such as text, tables, and images while extracting structural information together. Its key feature is going beyond simple text extraction to separate semantic units within documents and provide them as metadata.
In addition, HWP and HWPX documents provide paragraph-based hierarchical structure information, and PDF AI supports visual information interpretation functions based on Image Captioning.
(※ The image captioning function is currently in the PoC stage, and the commercial release schedule will be announced later.)
From Document Structure Analysis to Data Extraction
Hancom Data Loader generates structured data through a three-step processing flow: document input → document analysis → data extraction. The extracted data can then be linked with RAG pipelines and LLMs to build AI services.
In the document analysis stage, it runs two engines in parallel: rules-based analysis and AI-based document structure understanding. Structured documents with consistent formats are handled with rules, while unstructured documents that mix figures, tables, and images are analyzed precisely with AI.
It provides, in a single processing flow, the core capabilities required for document preprocessing—recognizing diverse document elements, detecting multi-column reading order, and extracting complex table and chart structures.
The extracted structured data can be directly integrated with Hancom’s proprietary RAG solution, HancomPedia, enabling rapid deployment of a document-based AI search system.

Document Data Extraction Targets: Support for Multiple Document Formats
Core Format Support for Public-Sector AI Adoption: HWP/HWPX Parsing
As the government moves to expand HWPX usage from 2026 and gradually restrict HWP attachments, full support for HWP and HWPX is becoming a prerequisite for public-sector AI adoption. Hancom Data Loader supports all HWP versions from 3.0 onward, and with a parser developed in-house by Hancom, it preserves a wide range of layout elements without loss without converting to PDF.
Universal Support for All Document Formats – PDF and OOXML
PDF AI handles structure analysis and table structure restoration through DLA, TSR, and OCR, and minimizes data loss in scanned documents through OCR integration. OOXML extracts text from office documents using a text extraction method.
Unstructured Visual Documents: PNG/JPG
PNG and JPG recognize text and structure to datafy even non-digitized materials such as fax documents, scanned images, and photographed documents.
👉Go to Hancom Data Loader Live Demo
Key Features of Hancom Data Loader: DLA/TSR/OCR-Based Document Structure Analysis
DLA – AI-Based Structure Analysis That Extracts Metadata
Hancom Data Loader’s DLA classifies objects into categories and extracts tables, footnotes, images, charts, captions, formulas, and even fonts as metadata.
When metadata is extracted together, RAG retrieval accuracy improves, enabling policy search, administrative QA, and internal knowledge agents to use highly reliable data.
Table/Image/Chart Structure Recognition: TSR Technology at the Core of Unstructured Data Processing
Key information in corporate documents is often concentrated in tables and charts rather than the main text. Based on TSR, Hancom Data Loader restores cell relationships—including borderless tables, merged cells, and tables within tables—to convert them into Markdown or HTML, and improves AI result quality by preserving multi-column layouts.
Data Separation and Post-Editing: Data Loader Studio
For documents where accuracy is difficult to ensure through automatic extraction alone, Data Loader Studio (an extended solution) is provided. Extraction results preprocessed by Data Loader can be reviewed side-by-side with the original document. Users can directly verify hierarchy, categories, reading order, and more, and post-correct necessary parts, enabling them to manage document structure and semantic relationships to fit actual business purposes.

Hancom Data Loader for document-based AI
🖥️Hancom Data Loader
Hancom Data Loader is a document parsing solution that converts HWP, HWPX, PDF, and OOXML into structured data. The extracted data can be integrated with Hancom’s RAG solution, Hancompedia, enabling an end-to-end workflow—from document collection to search and answers—built on a single Hancom stack. Hancom Data Loader provides a document preprocessing environment that agents can trust.
👉Go to Hancom Data Loader Live Demo
👉Inquire about Hancom Data Loader Implementation
References
1) Yonhap News, “”Reducing hwp that AI can’t read”… Government to transition to open hwpx,” 2026