Building a Digital Lexical Resource for Banyumasan Javanese: A Low-Resource Language Approach

  • Nisrina Hanifa Setiono Telkom University
  • Angga Kurniawan Telkom University
  • Viga Laksa Hardjanto Bina Nusantara University

Abstrak

Banyumasan Javanese, widely recognized through the Ngapak dialect, remains culturally significant but is still underrepresented in reusable computational resources. This study develops a digital Banyumasan-Indonesian lexical corpus and frames it as a reusable research artifact rather than a static appendix. The corpus was constructed from a Banyumasan-Indonesian dictionary, normalized into a structured bilingual dataset, and packaged as an installable Python resource so that it can be used directly in computational experiments. The implemented resource supports dataset loading, Banyumasan lookup, Indonesian lookup, simple translation, structured translation analysis, batch translation, and corpus statistics. The resulting corpus contains 2,000 lexical pairs, 1,996 unique Banyumasan forms, 1,444 unique Indonesian equivalents, and 4 duplicated Banyumasan headwords that preserve lexical ambiguity from the source material. To demonstrate practical utility, the study includes a 100-sentence implementation example in which Banyumasan text is translated with the published banyumasan-corpus package and evaluated against Indonesian ground truth using the Indonesian-focused embedding model LazarusNLP/all-indo-e5-small-v4. The average semantic similarity rises from 0.4833 for direct Banyumasan-versus-ground-truth comparison to 0.6427 after translation, producing an absolute gain of 0.1594 and a relative improvement of approximately 33.0% over the baseline. These findings indicate that a structured lexical corpus, when distributed in a directly reusable computational form, can strengthen both resource accessibility and small-scale downstream experimentation for a low-resource regional language.

Referensi

V. I. F. Maulidha and H. B. Thontowi, “The Influence of Familial Ethnic Socialization on Self-Esteem among Banyumasan Javanese Adolescents as Mediated by Ethnic Identity,” Jurnal Psikologi, vol. 52, no. 3, p. 295, 2025, doi: 10.22146/jpsi.109320.

C. Nugroho and I. P. Kusuma, “Identitas Budaya Banyumasan dalam Dialek Ngapak,” Jurnal Ilmu Komunikasi, vol. 21, no. 2, pp. 333–347, Sep. 2023, doi: 10.31315/JIK.V21I2.4556.

A. G. Pawestri, “Membangun Identitas Budaya Banyumasan Melalui Dialek Ngapak Di Media Sosial,” Jurnal Pendidikan Bahasa dan Sastra, vol. 19, no. 2, pp. 255–266, May 2020, doi: 10.17509/bs_jpbsp.v19i2.24791.

E. N. Purba, D. P. Togatorop, and A. Rosari, “Analisis Keterbatasan Korpus Bahasa dan Pemanfaatan AI dalam Penyusunan Kamus Batak Toba,” Jurnal Pendidikan Tambusai, vol. 9, no. 3, pp. 36479–36490, 2025.

A. Rokhman, R. E. Priyono, I. Santosa, S. Pangestuti, & Mustasyfa, and T. Kariadi, “Existence of Banyumasan Javanese Language in Digital Era,” Humanities and Social Science Research, vol. 5, no. 2, pp. p1–p1, May 2022, doi: 10.30560/HSSR.V5N2P1.

S. Cahyawijaya et al., “NusaCrowd: Open Source Initiative for Indonesian NLP Resources,” Proceedings of the Annual Meeting of the Association for Computational Linguistics, pp. 13745–13818, 2023, doi: 10.18653/V1/2023.FINDINGS-ACL.868.

H. Handoko, “Developing the Corpus of Minangkabau Language: Insights, Challenges, and Future Directions,” Jurnal Arbitrer, vol. 11, no. 3, pp. 413–429, Sep. 2024, doi: 10.25077/AR.11.3.413-429.2024.

B. K. Bal, B. Prasain, R. R. Ghimire, and P. Acharya, “Strategies for Corpus Development for Low-Resource Languages: Insights from Nepal,” Automatic Speech Recognition and Translation for Low Resource Languages, pp. 297–330, Jan. 2024, doi: 10.1002/9781394214624.CH15.

A. E. Nanda, V. P. Rantung, and K. Santa, “Development of a Web-Based Batak Simalungun Regional Language Corpus Using the Rapid Application Development Method,” Jurnal Teknik Informatika (Jutif), vol. 5, no. 4, pp. 549–558, 2024, doi: 10.52436/1.jutif.2024.5.4.2210.

D. I. Adelani et al., “MasakhaNER: Named Entity Recognition for African Languages,” Trans. Assoc. Comput. Linguist., vol. 9, pp. 1116–1131, Oct. 2021, doi: 10.1162/TACL_A_00416.

C. Pramartha and others, “Preserving the Enggano Language: A Digital Dictionary Approach,” ACIS Proceedings, pp. 1–8, 2024.

A. Tohari, P. Suratno, and A. Sudono, Kamus Bahasa Jawa Banyumasan Indonesia, Cetakan pe. Semarang: Balai Bahasa Provinsi Jawa Tengah, 2014.

S. Conggresco, V. P. Rantung, and Q. C. Kainde, “Development of a Web-Based Tonsea Language Corpus Using the Evolutionary Prototyping Method,” Jurnal Teknik Informatika (Jutif), vol. 5, no. 4, pp. 535–542, 2024, doi: 10.52436/1.jutif.2024.5.4.2201.

M. W. Goodman and F. Bond, “Intrinsically interlingual: The Wn Python library for wordnets,” GWC 2021 - Proceedings of the 11th Global Wordnet Conference, pp. 100–107, 2021, doi: 10.18653/v1/2021.gwc-1.12.

G. I. Winata et al., “NusaX: Multilingual Parallel Sentiment Dataset for 10 Indonesian Local Languages,” EACL 2023 - 17th Conference of the European Chapter of the Association for Computational Linguistics, Proceedings of the Conference, pp. 815–834, May 2022, doi: 10.18653/v1/2023.eacl-main.57.

J. M. List, R. Forkel, S. J. Greenhill, C. Rzymski, J. Englisch, and R. D. Gray, “Lexibank, a public repository of standardized wordlists with computed phonological and lexical features,” Scientific Data 2022 9:1, vol. 9, no. 1, pp. 316-, Jun. 2022, doi: 10.1038/s41597-022-01432-0.

Y. Prabowo, M. Gabriel, Nazarudin, T. Ratumanan, and M. Maslim, “Preserving Meher and Woirata Corpus Languages using Neural Machine Translation,” Indonesian Journal of Information Systems, vol. 6, no. 2, pp. 156–161, Feb. 2024, doi: 10.24002/IJIS.V6I2.8542.

A. F. Hidayatullah, R. A. Apong, D. T. C. Lai, and A. Qazi, “Word Level Language Identification in Indonesian-Javanese-English Code-Mixed Text,” Procedia Comput. Sci., vol. 244, pp. 105–112, Jan. 2024, doi: 10.1016/J.PROCS.2024.10.183.

K. Resiandi, Y. Murakami, and A. H. Nasution, “Neural Network-Based Bilingual Lexicon Induction for Indonesian Ethnic Languages,” Applied Sciences 2023, Vol. 13, Page 8666, vol. 13, no. 15, p. 8666, Jul. 2023, doi: 10.3390/APP13158666.

L. Wang, N. Yang, X. Huang, L. Yang, R. Majumder, and F. Wei, “Multilingual E5 Text Embeddings: A Technical Report,” Feb. 2024, Accessed: Apr. 16, 2026. [Online]. Available: https://arxiv.org/pdf/2402.05672

Diterbitkan
2026-07-24