{"id":483,"date":"2026-05-06T10:00:00","date_gmt":"2026-05-06T01:00:00","guid":{"rendered":"https:\/\/blog.hancom.com\/the-start-of-rag-llm-preprocessing-hancom-data-loader-document-parsing-solution-transforming-documents-into-ai-training-data\/"},"modified":"2026-07-16T15:16:51","modified_gmt":"2026-07-16T06:16:51","slug":"what-is-hancom-data-loader","status":"publish","type":"post","link":"https:\/\/blog.hancom.com\/en\/what-is-hancom-data-loader\/","title":{"rendered":"The Start of RAG\/LLM Preprocessing: Hancom Data Loader Document Parsing Solution &#8211; Transforming Documents into AI Training Data"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\"><\/p>\n\n<p class=\"wp-block-paragraph\"><strong><a href=\"https:\/\/sdk.hancom.com\/services\/1?type=DATA_LOADER\" target=\"_blank\" rel=\"noopener\">Hancom Data Loader<\/a> is a document parsing solution developed by Hancom.<\/strong> It is a document data preprocessing technology specialized for RAG-based AI training that helps AI understand HWP, HWPX, PDF, and OOXML.<\/p>\n\n<p class=\"wp-block-paragraph\">Before an agent can act, Hancom Data Loader converts documents into AI-readable data, separating and extracting content down to semantic units and providing it as agent-friendly metadata.<\/p>\n\n<div style=\"height:11px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n<h2 class=\"wp-block-heading\"><strong>Why is document datafication so difficult when adopting enterprise AI?<\/strong><\/h2>\n\n<p class=\"wp-block-paragraph\">Because large language models (LLMs) cannot read internal documents such as PDF, HWP, and HWPX as-is.<\/p>\n\n<p class=\"wp-block-paragraph\"> This unstructured document processing is the common barrier faced by public institutions, large-enterprise DX organizations, and SI partners considering in-house AI assistants, retrieval-augmented generation (RAG), and agent adoption.<\/p>\n\n<h3 class=\"wp-block-heading\"><strong><strong>Why Document Structuring Is Necessary: The Gap Between LLMs and Unstructured Documents<\/strong><\/strong><\/h3>\n\n<p class=\"wp-block-paragraph\">Enterprises hold vast amounts of materials\u2014manuals, reports, contracts, and official documents in HWP\/HWPX\u2014but not many are in a form that LLMs can use directly. <\/p>\n\n<p class=\"wp-block-paragraph\">In PDFs, table, multi-column, and caption structures disappear, and with HWP\/HWPX, meaning breaks with simple text extraction. This is why table rows and columns get mixed up and why answers can produce plausible summaries with missing evidence.  <strong>Only by converting documents into structured data can LLMs properly read internal information. <\/strong><\/p>\n\n<h3 class=\"wp-block-heading\"><strong><strong>Why Document Parsing Quality Matters in a RAG Pipeline<\/strong><\/strong><\/h3>\n\n<p class=\"wp-block-paragraph\">For organizations driving AI transformation (AX), structured document data is an asset that affects not only retrieval but also the reliability of agent decisions and execution. <\/p>\n\n<p class=\"wp-block-paragraph\">This is because document preprocessing quality in a RAG pipeline determines the entire process\u2014retrieval, reasoning, and execution. <\/p>\n\n<p class=\"wp-block-paragraph\">If table structures are lost during parsing, irrelevant text can be mixed in and mapped to the wrong locations during embedding.<\/p>\n\n<div style=\"height:10px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n<figure class=\"wp-block-image alignwide size-large\"><img decoding=\"async\" src=\"https:\/\/blog.hancom.com\/wp-content\/uploads\/2026\/05\/&#xACF5;&#xD1B5;-&#xD55C;&#xCEF4;&#xB370;&#xC774;&#xD130;&#xB85C;&#xB354;&#xB780;-1024x576.png\" alt=\"Hancom Data Loader, Developed by Hancom: An Automated Document Structure Analysis Solution\" class=\"wp-image-186\"\/><\/figure>\n\n<div style=\"height:10px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n<h2 class=\"wp-block-heading\"><strong><strong>Hancom Data Loader, a Document Structure Analysis Parsing Solution: Definition and Core Role<\/strong><\/strong><\/h2>\n\n<h3 class=\"wp-block-heading\"><strong>Hancom Data Loader: An Automated Document Structure Analysis Solution<\/strong><\/h3>\n\n<p class=\"wp-block-paragraph\"><strong><a href=\"https:\/\/sdk.hancom.com\/services\/1?type=DATA_LOADER\" data-type=\"link\" data-id=\"https:\/\/sdk.hancom.com\/services\/1?type=DATA_LOADER\" target=\"_blank\" rel=\"noopener\">Hancom Data Loader<\/a> is a document parsing solution developed by Hancom.<\/strong> It converts various document formats such as HWP, HWPX, PDF, and OOXML into structured data that AI can understand, and can be used to build RAG-based knowledge search systems, secure training data for Large Language Models (LLMs), and digitize corporate documents.<\/p>\n\n<p class=\"wp-block-paragraph\">It separates and extracts document semantic units\u2014including tables, titles, captions, and hierarchical structures\u2014based on OCR (Optical Character Recognition) extraction and provides them as agent-friendly metadata. As a document structure-preserving extraction method rather than general OCR, it is establishing itself as a core preprocessing technology for RAG construction. <br\/>Document analysis is performed based on Document Layout Analysis (DLA), OCR, and Table Structure Recognition (TSR). It distinguishes document components such as text, tables, and images, and extracts structural information together. Its key feature is that it goes beyond simple text extraction to separate semantic units within the document and provide them in metadata format. <\/p>\n\n<p class=\"wp-block-paragraph\">In addition, HWP and HWPX documents provide paragraph-based hierarchical structure information, and PDF_AI supports visual information interpretation functions based on Image Captioning.<\/p>\n\n<p class=\"wp-block-paragraph\">(\u203b The image captioning function is currently in the PoC stage, and the commercial release schedule will be announced later.)<\/p>\n\n<div style=\"height:10px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n<h3 class=\"wp-block-heading\"><strong>From Document Structure Analysis to Data Extraction<\/strong><\/h3>\n\n<p class=\"wp-block-paragraph\">Hancom Data Loader generates structured data through a three-step processing flow: <strong>document input \u2192 document analysis \u2192 data extraction<\/strong>. The extracted data can then be linked with RAG pipelines and LLMs to build AI services. <\/p>\n\n<p class=\"wp-block-paragraph\">In the document analysis stage, it runs two engines in parallel: rules-based analysis and AI-based document structure understanding. Structured documents with consistent formats are handled with rules, while unstructured documents that mix figures, tables, and images are analyzed precisely with AI.  <\/p>\n\n<p class=\"wp-block-paragraph\">It provides, in a single processing flow, the core capabilities required for document preprocessing\u2014recognizing diverse document elements, detecting multi-column reading order, and extracting complex table and chart structures.<\/p>\n\n<p class=\"wp-block-paragraph\"><strong>The extracted structured data can be directly integrated with Hancom\u2019s proprietary RAG solution, HancomPedia, enabling rapid deployment of a document-based AI search system. <\/strong><\/p>\n\n<div style=\"height:10px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n<figure class=\"wp-block-image alignwide size-large\"><img decoding=\"async\" src=\"https:\/\/blog.hancom.com\/wp-content\/uploads\/2026\/05\/&#xBE14;&#xB85C;&#xADF8;-&#xCF58;&#xD150;&#xCE20;-01-1024x576.png\" alt=\"Image explaining the AI document processing and data utilization process\" class=\"wp-image-67\"\/><figcaption class=\"wp-element-caption\">A four-step pipeline combining Hancom Data Loader and HancomPedia, processing the entire workflow from document input to utilization within a single solution<\/figcaption><\/figure>\n\n<div style=\"height:10px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n<h3 class=\"wp-block-heading\"><strong>Document Data Extraction Targets: Support for Multiple Document Formats<\/strong><\/h3>\n\n<h4 class=\"wp-block-heading\"><strong><strong>Core Format Support for Public-Sector AI Adoption: HWP\/HWPX Parsing<\/strong><\/strong><\/h4>\n\n<p class=\"wp-block-paragraph\">As the government moves to expand HWPX usage from 2026 and <a href=\"https:\/\/www.yna.co.kr\/view\/AKR20260423082600017\" target=\"_blank\" rel=\"noopener\">gradually restrict HWP attachments<\/a>, <strong>full support for HWP and HWPX is becoming a prerequisite for public-sector AI adoption<\/strong>. Hancom Data Loader supports all HWP versions from 3.0 onward, and with a parser developed in-house by Hancom, it preserves more than 20 layout elements <strong>without loss<\/strong> <strong>without converting to PDF<\/strong>.<\/p>\n\n<h4 class=\"wp-block-heading\"><strong>Universal Format Support: PDF and OOXML (DOCX, XLSX, PPTX)<\/strong><\/h4>\n\n<p class=\"wp-block-paragraph\">PDF(AI) handles layout analysis and table structure restoration through Document Layout Analysis (DLA), Table Structure Recognition (TSR), and OCR, and minimizes data loss for scanned documents through OCR integration. OOXML (DOCX, XLSX, PPTX) extracts text from Office documents using a text extraction method. <\/p>\n\n<h4 class=\"wp-block-heading\"><strong>Unstructured Visual Documents: PNG\/JPG<\/strong><\/h4>\n\n<p class=\"wp-block-paragraph\">PNG and JPG automatically recognize text and layout to datafy even non-digitized materials such as fax documents, scanned images, and photographed documents.<\/p>\n\n<p class=\"wp-block-paragraph\"><strong>\ud83d\udc49<\/strong><a href=\"https:\/\/livedemo.sdk.hancom.com\/dataloader\" target=\"_blank\" rel=\"noopener\"><strong>Go to Hancom Data Loader Live Demo<\/strong><\/a><\/p>\n\n<div style=\"height:10px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n<h2 class=\"wp-block-heading\"><strong><strong><strong>Hancom Data Loader Key Features &#8211; Document Structure Analysis based on DLA, TSR, and OCR<\/strong><\/strong><\/strong><\/h2>\n\n<h3 class=\"wp-block-heading\"><strong><strong><strong>Document Layout Analysis (DLA) &#8211; AI-based structural analysis that extracts even metadata<\/strong><\/strong><\/strong><\/h3>\n\n<p class=\"wp-block-paragraph\">Hancom Data Loader&#8217;s Document Layout Analysis (DLA) classifies objects into categories and extracts tables, footnotes, images, charts, captions, formulas, and even fonts as metadata.<\/p>\n\n<p class=\"wp-block-paragraph\">When metadata is extracted together, RAG retrieval accuracy improves, enabling policy search, administrative QA, and internal knowledge agents to use highly reliable data.<\/p>\n\n<h3 class=\"wp-block-heading\"><strong><strong><strong>Table, Image, and Chart Structure Recognition &#8211; Table Structure Recognition (TSR) technology, the core of unstructured data processing<\/strong><\/strong><\/strong><\/h3>\n\n<p class=\"wp-block-paragraph\">Key information in corporate documents is often concentrated in tables and charts rather than the main text. Based on Table Structure Recognition (TSR), Hancom Data Loader restores cell relationships\u2014including borderless tables, merged cells, and tables within tables\u2014to convert them into Markdown or HTML, and improves AI result quality by preserving multi-column layouts. <\/p>\n\n<h3 class=\"wp-block-heading\"><strong>Data Separation and Post-Editing: Data Loader Studio<\/strong><\/h3>\n\n<p class=\"wp-block-paragraph\">We provide Data Loader Studio (an extension solution) for documents where it is difficult to ensure accuracy through automatic extraction alone. Extraction results preprocessed with Data Loader can be reviewed by comparing them side-by-side with the original document. Users can directly check the hierarchy, category, and reading order, and post-edit necessary parts to ensure that the document structure and semantic relationships are refined to suit actual business purposes. In addition, through semantic-based tagging and labeling, extraction criteria can be supplemented to fit the customer&#8217;s document format and refined into data quality suitable for training and searching.  <\/p>\n\n<div style=\"height:10px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n<figure class=\"wp-block-image alignwide size-large\"><img decoding=\"async\" src=\"https:\/\/blog.hancom.com\/wp-content\/uploads\/2026\/05\/&#xB370;&#xC774;&#xD130;-&#xB85C;&#xB354;-&#xD65C;&#xC6A9;-&#xD50C;&#xB85C;&#xC6B0;-1024x576.png\" alt=\"Image explaining the Hancom Data Loader document utilization flow \" class=\"wp-image-135\"\/><\/figure>\n\n<div style=\"height:10px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n<h2 class=\"wp-block-heading\"><strong>The Starting Point for Document-Based AI: Hancom Data Loader<\/strong><\/h2>\n\n<h3 class=\"wp-block-heading\"><a href=\"https:\/\/sdk.hancom.com\/services\/1?type=DATA_LOADER\" data-type=\"link\" data-id=\"https:\/\/sdk.hancom.com\/services\/1?type=DATA_LOADER\" target=\"_blank\" rel=\"noopener\"><strong>\ud83d\udda5\ufe0f Hancom Data Loader<\/strong><\/a><\/h3>\n\n<p class=\"wp-block-paragraph\"><strong><a href=\"https:\/\/blog.hancom.com\/en\/hancom-data-loader-startguide\/\">Hancom Data Loader<\/a> is a document parsing solution that converts HWP, HWPX, PDF, and OOXML into structured data<\/strong>. The extracted data can be integrated with our RAG solution, HancomPedia, enabling a single Hancom stack from document collection to search and answers. Start with document preprocessing that agents can trust\u2014start with Hancom Data Loader.  <\/p>\n\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/livedemo.sdk.hancom.com\/dataloader\" target=\"_blank\" rel=\"noopener\"><strong>\ud83d\udc49 Go to Hancom Data Loader Live Demo<\/strong><\/a><\/p>\n\n<p class=\"wp-block-paragraph\"><strong>\ud83d\udc49<\/strong><a href=\"https:\/\/sdk.hancom.com\/contacts\/create\" target=\"_blank\" rel=\"noopener\"><strong>Inquire about Hancom Data Loader Implementation<\/strong><\/a><\/p>\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n<p class=\"wp-block-paragraph\"><strong>References<\/strong><\/p>\n\n<p class=\"wp-block-paragraph\">1) <a href=\"https:\/\/www.yna.co.kr\/view\/AKR20260423082600017\" data-type=\"link\" data-id=\"https:\/\/www.yna.co.kr\/view\/AKR20260423082600017\" target=\"_blank\" rel=\"noopener\"><strong>Yonhap News<\/strong><\/a>, \u201c&#8221;Reducing hwp that AI can\u2019t read&#8221;\u2026 Government to transition to open hwpx,\u201d 2026<\/p>\n\n<p class=\"wp-block-paragraph\"><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Hancom Data Loader is a document data preprocessing technology specialized for RAG-based AI training that helps AI understand HWP, HWPX, PDF, and OOXML. We have compiled everything you need to know about Data Loader, a core preprocessing technology for building RAG solutions and implementing corporate AI. <\/p>\n","protected":false},"author":2,"featured_media":486,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[30],"tags":[],"class_list":["post-483","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-hancom-product-user-guide"],"_links":{"self":[{"href":"https:\/\/blog.hancom.com\/en\/wp-json\/wp\/v2\/posts\/483","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blog.hancom.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blog.hancom.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blog.hancom.com\/en\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/blog.hancom.com\/en\/wp-json\/wp\/v2\/comments?post=483"}],"version-history":[{"count":8,"href":"https:\/\/blog.hancom.com\/en\/wp-json\/wp\/v2\/posts\/483\/revisions"}],"predecessor-version":[{"id":904,"href":"https:\/\/blog.hancom.com\/en\/wp-json\/wp\/v2\/posts\/483\/revisions\/904"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/blog.hancom.com\/en\/wp-json\/wp\/v2\/media\/486"}],"wp:attachment":[{"href":"https:\/\/blog.hancom.com\/en\/wp-json\/wp\/v2\/media?parent=483"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blog.hancom.com\/en\/wp-json\/wp\/v2\/categories?post=483"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blog.hancom.com\/en\/wp-json\/wp\/v2\/tags?post=483"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}