全部 标题 作者
关键词 摘要

OALib Journal期刊
ISSN: 2333-9721
费用:99美元

查看量下载量

Corpus Construction and NER Model Comparison for Formation Name Recognition in Completion Reports

DOI: 10.4236/oalib.1115785, PP. 1-8

Subject Areas: Linguistics

Keywords: Completion Reports, Oil and Gas, Corpus Annotation, Conditional Random Fields

Full-Text   Cite this paper   Add to My Lib

Abstract

Completion reports are central to the oil and gas exploration and development process, in which formation names serve as fundamental entities for geological modeling and reserve calculation. However, general-purpose Chinese named entity recognition (NER) tools do not include formation names as an entity type, and the petroleum domain lacks Chinese corpus annotation at the sequence labeling level. Based on the World Oil Outlook 2026 (WOO 2026) published by OPEC, this study employs AntConc frequency analysis and KH Coder co-occurrence network analysis to design data-driven annotation rules, and completes independent annotation and adjudication review through Doccano. The results show that the fully-featured conditional random field (CRF) model combined with rule-based post-processing achieves an F1 score significantly superior to the rule-based baseline. This provides foundational resources and methodological references for Chinese information extraction in the petroleum domain.

Cite this paper

Cai, Y. (2026). Corpus Construction and NER Model Comparison for Formation Name Recognition in Completion Reports. Open Access Library Journal, 13, e15785. doi: http://dx.doi.org/10.4236/oalib.1115785.

References

[1]  OPEC (2026) World Oil Outlook 2026. OPEC.
[2]  Lafferty, J., McCallum, A. and Pereira, F.C.N. (2001) Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data. <i>Proceedings of the </i>18<i>th International Conference on Machine Learning</i>, San Francisco, 28 June 2001-1 July 2001 282-289.
[3]  Huang, Z., Xu, W. and Yu, K. (2015) Bidirectional LSTM-CRF Models for Sequence Tagging. arXiv:1508.01991.
[4]  Devlin, J., Chang, M.W., Lee, K., <i>et al</i><i>.</i> (2019) BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding. <i>Proceedings of NAACL</i>-<i>HL</i><i>T</i>, Minneapolis, June 2-June 7 2019, 4171-4186. <br>https://aclanthology.org/N19-1423/
[5]  Lee, J., Yoon, W., Kim, S., Kim, D., Kim, S., So, C.H., <i>et al</i>. (2019) BioBERT: A Pre-Trained Biomedical Language Representation Model for Biomedical Text Mining. <i>Bioinformatics</i>, 36, 1234-1240. <br>https://doi.org/10.1093/bioinformatics/btz682
[6]  Leitner, E., Rehm, G. and Moreno-Schneider, J. (2019) Fine-Grained Named Entity Recognition in Legal Documents. In: Acosta, M., Cudr&#233;-Mauroux, P., Maleshkova, M., Pellegrini, T., Sack, H. and Sure-Vetter, Y., Eds., <i>Lecture</i> <i>Notes</i> <i>in</i> <i>Computer</i> <i>Science</i>, Springer International Publishing, 272-287. <br>https://doi.org/10.1007/978-3-030-33220-4_20
[7]  Enkhsaikhan, M., Holden, E., Duuring, P. and Liu, W. (2021) Understanding Ore-Forming Conditions Using Machine Reading of Text. <i>Ore Geology Reviews</i>, 135, 104200. <br>https://doi.org/10.1016/j.oregeorev.2021.104200

Full-Text


Contact Us

service@oalib.com

QQ:3279437679

WhatsApp +8615387084133