Completion reports are central to the oil and gas exploration and development process, in which formation names serve as fundamental entities for geological modeling and reserve calculation. However, general-purpose Chinese named entity recognition (NER) tools do not include formation names as an entity type, and the petroleum domain lacks Chinese corpus annotation at the sequence labeling level. Based on the World Oil Outlook 2026 (WOO 2026) published by OPEC, this study employs AntConc frequency analysis and KH Coder co-occurrence network analysis to design data-driven annotation rules, and completes independent annotation and adjudication review through Doccano. The results show that the fully-featured conditional random field (CRF) model combined with rule-based post-processing achieves an F1 score significantly superior to the rule-based baseline. This provides foundational resources and methodological references for Chinese information extraction in the petroleum domain.
Cite this paper
Cai, Y. (2026). Corpus Construction and NER Model Comparison for Formation Name Recognition in Completion Reports. Open Access Library Journal, 13, e15785. doi: http://dx.doi.org/10.4236/oalib.1115785.
Lafferty, J., McCallum, A. and Pereira, F.C.N. (2001) Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data. <i>Proceedings of the </i>18<i>th International Conference on Machine Learning</i>, San Francisco, 28 June 2001-1 July 2001 282-289.
Devlin, J., Chang, M.W., Lee, K., <i>et al</i><i>.</i> (2019) BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding. <i>Proceedings of NAACL</i>-<i>HL</i><i>T</i>, Minneapolis, June 2-June 7 2019, 4171-4186. <br>https://aclanthology.org/N19-1423/
Lee, J., Yoon, W., Kim, S., Kim, D., Kim, S., So, C.H., <i>et al</i>. (2019) BioBERT: A Pre-Trained Biomedical Language Representation Model for Biomedical Text Mining. <i>Bioinformatics</i>, 36, 1234-1240. <br>https://doi.org/10.1093/bioinformatics/btz682
Leitner, E., Rehm, G. and Moreno-Schneider, J. (2019) Fine-Grained Named Entity Recognition in Legal Documents. In: Acosta, M., Cudré-Mauroux, P., Maleshkova, M., Pellegrini, T., Sack, H. and Sure-Vetter, Y., Eds., <i>Lecture</i> <i>Notes</i> <i>in</i> <i>Computer</i> <i>Science</i>, Springer International Publishing, 272-287. <br>https://doi.org/10.1007/978-3-030-33220-4_20
Enkhsaikhan, M., Holden, E., Duuring, P. and Liu, W. (2021) Understanding Ore-Forming Conditions Using Machine Reading of Text. <i>Ore Geology Reviews</i>, 135, 104200. <br>https://doi.org/10.1016/j.oregeorev.2021.104200