Optimasi LayoutLMv3 untuk DocVQA pada Tabel Statistik Melalui Fine-Tuning dan Semantic Mapping
Optimization of LayoutLMv3 for DocVQA on Statistical Tables via Fine-Tuning and Semantic Mapping
DOI:
https://doi.org/10.57152/malcom.v6i3.2632Keywords:
Continual Fine-Tuning, DocVQA, LayoutLMv3, Optimasi Model, Tabel StatistikAbstract
Penelitian ini mengoptimalkan model LayoutLMv3 untuk tugas Document Visual Question Answering (DocVQA) pada tabel statistik melalui kombinasi Continual Fine-tuning (CFT) yang mengintegrasikan normalisasi koordinat 2D dan patch embedding visual, serta lapisan logika Semantic Similarity Mapping (SSM) untuk menangani variasi pertanyaan. Dataset mandiri dibangun dari 398 gambar tabel publikasi BPS "Kabupaten Pinrang Dalam Angka 2025" dengan 19.900 pasangan QA. Hasil evaluasi menunjukkan Training Loss turun dari 3,72 menjadi 0,63, ROUGE-1 mencapai 74,06%, dan akurasi human validation meningkat dari 0% (zero-shot) menjadi 41% (CFT) dan 51% (CFT+SSM). Akurasi tertinggi 100% dicapai pada tabel dengan header tiga level. Pendekatan multimodal ini efektif meningkatkan pemahaman struktur tabel statistik, namun generalisasi pada tabel kompleks masih terbatas oleh dataset.
Downloads
References
Y. Huang, T. Lv, L. Cui, Y. Lu, and F. Wei, "LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking," in Proceedings of the 30th ACM International Conference on Multimedia (MM '22), Association for Computing Machinery, Oct. 2022, pp. 4083–4091. doi: 10.1145/3503161.3548112.
K. Kapula, "Intelligent Document Processing: The New Frontier of Automation," International Journal of Intelligent Systems and Applications in Engineering, vol. 10, no. 2, pp. 215–223, 2022. doi: 10.18201/ijisae.2022.272.
M. Mathew, D. Karatzas, and C. V. Jawahar, "DocVQA: A Dataset for VQA on Document Images," in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), IEEE, Jan. 2021, pp. 2200–2209. doi: 10.1109/WACV48630.2021.00225.
N. D. Huynh, M. R. Bouadjenek, S. Aryal, I. Razzak, and H. Hacid, "Visual Question Answering: From Early Developments to Recent Advances A Survey," arXiv preprint, arXiv:2501.03939, Jan. 2025. [Online]. Available: http://arxiv.org/abs/2501.03939.
T. Sumner and M. Marlino, "Digital Libraries and Educational Practice: Opportunities and Challenges," in Proceedings of the 4th ACM/IEEE-CS Joint Conference on Digital Libraries (JCDL '04), Association for Computing Machinery, Jun. 2004, pp. 170–178. doi: 10.1145/996350.996389.
B. S. U. Kim, J. Kim, D. Lee, and B. Jang, "Visual Question Answering: A Survey of Methods, Datasets, Evaluation, and Challenges," ACM Computing Surveys, vol. 57, no. 10, pp. 1–42, May 2025. doi: 10.1145/3728635.
R. Tanaka, K. Nishida, K. Nishida, T. Hasegawa, I. Saito, and K. Saito, "SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images," in Proceedings of the 37th AAAI Conference on Artificial Intelligence (AAAI-23), AAAI Press, Feb. 2023, pp. 13636–13645. doi: 10.1609/aaai.v37i11.26598.
C. Barboule, B. Piwowarski, and Y. Chabot, "Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends," arXiv preprint, arXiv:2501.02235, Mar. 2025. [Online]. Available: http://arxiv.org/abs/2501.02235.
Q. Peng, Y. Tang, Y. Xu, H. Zhao, G. Lv, B. Sun, Y. Lv, and Y. Wei, "ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding," in Findings of the Association for Computational Linguistics: EMNLP 2022, Association for Computational Linguistics, Dec. 2022, pp. 3744–3756. doi: 10.18653/v1/2022.findings-emnlp.274.
S. Chanti, P. Harika, V. Venkata Vinay, P. Guna Sri Charan, and Sk. Ajad, "DocumentQA: Leveraging LayoutLMv3 for Next-Level Question Answering," in Proceedings of the 2023 International Conference on Self Sustainable Artificial Intelligence Systems (ICSSAS), IEEE, Oct. 2023, pp. 1–8. doi: 10.1109/ICSSAS57918.2023.10331881.
M. Yang, A. Anantharaman, Z. Kitowski, and D. C. Robert, "Graph Relation Transformer: Incorporating Pairwise Object Features into the Transformer Architecture," arXiv preprint, arXiv:2111.06075, Nov. 2021. [Online]. Available: http://arxiv.org/abs/2111.06075.
J. Park, J. Cho, S. Han, and K. Kim, "OCR-Augmented GPT for Accurate Text Extraction in Industrial Environments," IEEE Access, vol. 13, pp. 136226–136236, 2025. doi: 10.1109/ACCESS.2025.3594682.
G. A. Pereira and M. Hussain, "A Review of Transformer-Based Models for Computer Vision Tasks: Capturing Global Context and Spatial Relationships," arXiv preprint, arXiv:2408.15178, Aug. 2024. [Online]. Available: http://arxiv.org/abs/2408.15178.
A. R. I. Ponggohong, G. C. Rorimpandey, and V. V. Maswonggo, "Penerapan Optical Character Recognition Tesseract dan Gemini pada Sistem Buku Tamu Digital Berbasis Web," Jurnal Informatika dan Teknologi Komputer, vol. 6, no. 1, pp. 45–53, 2025. doi: 10.46808/jitkom.v6i1.312.
S. Gull, N. Ahmed, and S. A. Sheikh, "Deep learning-Driven OCR System for Brahui Printed Text: Bridging the Digital Gap in Low-Resource Language Processing," Liberal Journal of Language and Literature Review, vol. 3, no. 1, pp. 88–102, 2025. [Online]. Available: https://llrjournal.com/index.php/11.
A. Auriemma Citarella, M. Barbella, M. G. Ciobanu, F. De Marco, L. Di Biasi, and G. Tortora, "Assessing the Effectiveness of ROUGE as Unbiased Metric in Extractive vs. Abstractive Summarization Techniques," Journal of Computational Science, vol. 87, art. no. 102571, May 2025. doi: 10.1016/j.jocs.2025.102571.
M. Gabburo, S. Garg, R. Koncel-Kedziorski, and A. Moschitti, "SQUARE: Automatic Question Answering Evaluation using Multiple Positive and Negative References," in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL 2023), Association for Computational Linguistics, Jul. 2023, pp. 441–452. doi: 10.18653/v1/2023.acl-short.49.
L. M. Amugongo, P. Mascheroni, S. Brooks, S. Doering, and J. Seidel, "Retrieval Augmented Generation for Large Language Models in Healthcare: A Systematic Review," PLOS Digital Health, vol. 4, no. 6, art. no. e0000877, Jun. 2025. doi: 10.1371/journal.pdig.0000877.
BPS Kabupaten Pinrang, Kabupaten Pinrang Dalam Angka 2025, vol. XXI. Pinrang: BPS Kabupaten Pinrang, 2025. [Online]. Available: https://pinrangkab.bps.go.id.
T. Perdana, K. Kusnandar, H. H. Perdana, F. R. Hermiatin, A. H. Sadeli, and R. Saville, "Behavioural Drivers of Thiamethoxam Use among Perennial and Annual Crop Farmers in Indonesia: An Extended Theory of Planned Behavior Approach," Cogent Food & Agriculture, vol. 12, no. 1, art. no. 2615202, Dec. 2026. doi: 10.1080/23311932.2026.2615202.
Y. Xu, M. Xu, F. Lv, L. Cui, G. Wei, N. Wang, F. Lu, B. Florencio, C. Zhang, W. Che, M. Zhang, and L. Zhou, "LayoutLMv2: Multi-modal Pre-training for Visually-rich Document Understanding," in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-IJCNLP 2021), Association for Computational Linguistics, Aug. 2021, pp. 2579–2591. doi: 10.18653/v1/2021.acl-long.201.
D. Caffagni, S. Sarto, M. Cornia, L. Baraldi, and R. Cucchiara, "Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document Retrieval," in Proceedings of the 2022 ACM International Conference on Multimedia Retrieval (ICMR '22), Association for Computing Machinery, Jun. 2022, pp. 180–188. doi: 10.1145/3512527.3531391.
C. Gao, Q. Zeng, M. Yin, and J. Sun, "Structured Multimodal Attentions for TextVQA," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 2, pp. 1594–1607, Feb. 2023. doi: 10.1109/TPAMI.2022.3152766.
N. Hegde, S. Paul, G. Madan, and G. Aggarwal, "Analyzing the Efficacy of an LLM-Only Approach for Image-based Document Question Answering," arXiv preprint, arXiv:2309.14389, Sep. 2023. [Online]. Available: http://arxiv.org/abs/2309.14389.
R. Pal, S. Kar, D. K. Prasad, and A. A. Sekh, "VisionAidQA: Advancing Visual Question Answering for the Visually Impaired," ACM Computing Surveys, vol. 57, no. 3, art. no. 3788861, Jul. 2026. doi: 10.1145/3788861.
X. Wu, W. Hu, X. Hu, and E. Chang, "A Region-based Document VQA," in Proceedings of the 30th ACM International Conference on Multimedia (MM '22), Association for Computing Machinery, Oct. 2022, pp. 4909–4920. doi: 10.1145/3503161.3548172.
J. Nockels, P. Gooding, S. Ames, and M. Terras, "Understanding the Application of Handwritten Text Recognition Technology in Heritage Contexts: A Systematic Review of Transkribus in Published Research," Archival Science, vol. 22, no. 3, pp. 367–392, Sep. 2022. doi: 10.1007/s10502-022-09397-0.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 D. Achmad Ansari A., Hazriani, Yuyun Yuyun

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Copyright © by Author; Published by Institut Riset dan Publikasi Indonesia (IRPI)
This Indonesian Journal of Machine Learning and Computer Science is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.










