oalib

OALib Journal期刊

ISSN: 2333-9721

费用:99美元

投稿

时间不限

( 2026 )

( 2025 )

( 2024 )

( 2023 )

自定义范围…

匹配条件: “Part-of-speech” ,找到相关结果约1000条。
列表显示的所有文章,均可免费获取
第1页/共1000条
每页显示 条
基于依存句法与词性标注的汉英翻译分析与研究
A Study on Chinese-English Translation Based on Dependency Syntax and Part-of-Speech Tagging
 [PDF]

陈佳仪
Modern Linguistics (ML) , 2026, DOI: 10.12677/ml.2026.141063
Abstract: 本论文通过构建语料库,利用自然语言处理工具spaCy对语料进行词性标注和依存关系分析,探讨汉译英翻译中的语言差异。研究表明汉语更倾向使用主动语态和名词短语,而英语更依赖被动语态和介词短语表达复杂概念,为翻译实践提供了新的见解,也助于优化翻译策略及机器翻译系统。
By constructing a corpus and using spaCy, a natural language processing tool, to conduct part-of-speech tagging and dependency relation analysis on the corpus, this paper explores the linguistic differences in Chinese-English translation. The research shows that Chinese tends to use the active voice and noun phrases, whereas English relies more on the passive voice and prepositional phrases to express complex concepts. This study provides new insights for translation practice and also helps to optimize translation strategies and machine translation systems.
Sentence Clustering Using Parts-of-Speech
Richard Khoury
International Journal of Information Engineering and Electronic Business , 2012,
Abstract: Clustering algorithms are used in many Natural Language Processing (NLP) tasks. They have proven to be popular and effective tools to use to discover groups of similar linguistic items. In this exploratory paper, we propose a new clustering algorithm to automatically cluster together similar sentences based on the sentences’ part-of-speech syntax. The algorithm generates and merges together the clusters using a syntactic similarity metric based on a hierarchical organization of the parts-of-speech. We demonstrate the features of this algorithm by implementing it in a question type classification system, in order to determine the positive or negative impact of different changes to the algorithm.
Towards the Construction of a Gold Standard Biomedical Corpus for the Romanian Language
Grigorina Mitrofan,Maria Mitrofan,Verginica Barbu Mititelu
Data | An Open Access Journal from MDPI , 2018, DOI: https://doi.org/10.3390/data3040053
Abstract: Abstract Gold standard corpora (GSCs) are essential for the supervised training and evaluation of systems that perform natural language processing (NLP) tasks. Currently, most of the resources used in biomedical NLP tasks are mainly in English. Little effort has been reported for other languages including Romanian and, thus, access to such language resources is poor. In this paper, we present the construction of the first morphologically and terminologically annotated biomedical corpus of the Romanian language (MoNERo), meant to serve as a gold standard for biomedical part-of-speech (POS) tagging and biomedical named entity recognition (bioNER). It contains 14,012 tokens distributed in three medical subdomains: cardiology, diabetes and endocrinology, extracted from books, journals and blogposts. In order to automatically annotate the corpus with POS tags, we used a Romanian tag set which has 715 labels, while diseases, anatomy, procedures and chemicals and drugs labels were manually annotated for bioNER with a Cohen Kappa coefficient of 92.8% and revealed the occurrence of 1877 medical named entities. The automatic annotation of the corpus has been manually checked. The corpus is publicly available and can be used to facilitate the development of NLP algorithms for the Romanian language. View Full-Tex
兼类词概率分布计量考察及语法搭配模式在中文信息处理中的应用
A Study of the Probability Distribution and Grammatical Collocation Patterns of Multi-Category Words in Chinese Information Processing
 [PDF]

王浩学, 徐艳华
Modern Linguistics (ML) , 2021, DOI: 10.12677/ML.2021.92072
Abstract: 在词性标注的过程中,汉语中兼类词的存在是影响词性标注准确率的主要原因。本研究以三部词典标注一致的78个形名兼类词为测试对象,基于规则和统计相结合的词性标注方法,将统计的兼类词分布概率与语法搭配规则结合起来,利用兼类词语法搭配模式构建规则库,对国家语委现代汉语通用平衡语料库标注的兼类词结果进行修正,准确率可以提高14.57%。
In the process of part-of-speech tagging, the existence of multi-category words in Chinese is the main reason that affects the accuracy of part-of-speech tagging. In this study, 78 adjective-noun multi-category words of the same part-of-speech tagging in the three dictionaries are the test objects. The part-of-speech tagging method based on the combination of rules and statistics combines the statistical distribution probability of multi-category words with grammatical collocation rules, and builds a rule database using the grammatical collocation mode of multi-category words. The rule database corrects the results of the multi-category words tagged by the modern Chinese corpus of State Language Commission, and the accuracy rate can be increased by 14.57%.
词典中“操作”的词类标注问题——一项基于语料库的个案研究
The Problem of Part-of-Speech Labeling for “Cao Zuo” in Dictionaries—A Case Study Based on Corpus
 [PDF]

郝晨, 唐菊, 孟波
Modern Linguistics (ML) , 2023, DOI: 10.12677/ML.2023.1110626
Abstract: 词类标注是词典的一项重要任务,词类标注的准确性直接影响词典的编纂质量。典型名词和典型动词的标注较容易进行,但对于兼类词的标注争议较大,兼类词的词类标注一直是词典编纂的难点。本研究以语料库作为研究工具,选取词类识别有争议的“操作”一词进行个案研究,考察现有词典对“操作”一词的释义情况,初步判断其词类。同时结合语料库中“操作”一词的使用模式,通过搭配和类联接等方法进一步确定“操作”一词的词性处理,试图对“操作”一词进行准确的词类标注。研究发现“操作”在现代汉语中应同时属于名词、动词,词典应将其处理为名动兼类词。
Part-of-speech labeling is an important task in dictionaries, and the accuracy of part-of-speech la-beling directly affects the quality of dictionary compilation. The labeling of typical nouns and verbs is relatively easy, but there is a significant controversy over the labeling of concurrent words. The labeling of concurrent words has always been a difficulty in dictionary compilation. This study uses a corpus as a research tool and selects the controversial word “Cao Zuo” for word class recognition as a case study. It examines the interpretation of the word “Cao Zuo” in existing dictionaries and pre-liminarily determines its word class. At the same time, combining the usage patterns of the word “Cao Zuo” in the corpus, further determining the part-of-speech processing of the word “Cao Zuo” through methods such as collocation and class connections, attempting to accurately label the part of speech of the word “Cao Zuo”. Research has found that “Cao Zuo” should belong to both nouns and verbs in Chinese, and dictionaries should treat it as a noun-verb hybrid word.
Análise Morfossintáctica para Português Europeu e Galego: Problemas, Solu es e Avalia o
Marcos Garcia,Pablo Gamallo
Linguamática , 2010,
Abstract: As diferentes tarefas de análise morfossintáctica têm muita importancia para posteriores níveis do processamento da linguagem natural. Por isso, estes processos devem ser realizados com ferramentas que garantam bons desempenhos em rela o à cobertura, precis o e robustez na análise. FreeLing é uma suite com licen a GPL desenvolvida pelo Grupo TALP da Universitat Politècnica de Catalunya. Este software contém -entre outros- módulos de tokeniza o, segmenta o de ora es, reconhecimento de entidades e anota o morfossintáctica. Com o fim de obtermos ferramentas que nos sirvam de base para a análise sintáctica, bem como para disponibilizar software livre para o processamento de superfície de Português Europeu e Galego, adaptámos FreeLing para estas variedades. A primeira delas foi desenvolvida com ajuda de recursos linguísticos disponíveis on-line, enquanto os ficheiros do Galego tiveram como base a vers o anterior de FreeLing (criados pelo Seminario de Lingüística Informática da Universidade de Vigo), que já realizava a análise desta língua. O presente trabalho descreve os principais aspectos do desenvolvimento das ferramentas, com ênfase nos problemas encontrados e nas solu es adoptadas em cada caso. Além disso, s o apresentados os resultados de avalia o do módulo PoS-tagger.
Integral theory of polysemy
Sternina M.A.
Ju?noslovenski Filolog , 2006, DOI: 10.2298/jfi0662224s
Abstract: Integral theory of polysemy is agued which represents the most general view on the problem of polysemy. The notion of polysemy is essentially extended and is applied to both lexical and grammatical language levels. It is argued that polysemy regulates and systematizes both vocabulary and grammar and may be considered as a factor which is organizing the language system.
More Effective Web Search Using Bigrams and Trigrams
David Johnson,Vishv Malhotra,Peter Vamplew
Webology , 2006,
Abstract: This paper investigates the effectiveness of quoted bigrams and trigrams as query terms to target web search. Prior research in this area has largely focused on static corpora each containing only a few million documents, and has reported mixed (usually negative) results. We investigate the bigram/trigram extraction problem and present an extraction algorithm that shows promising results when applied to real-time web search. We also present a prototype augmented search software package that can leverage the results provided by a web search engine to assist the web searcher identify important phrases and related documents quickly. This software has received favourable feedback in a recent user survey.
A Tolerance Rough Set Based Semantic Clustering Method for Web Search Results
Xian-Jun Meng,Qing-Cai Chen,Xiao-Long Wang
Information Technology Journal , 2009,
Abstract: The objective of this study is to present a new web search results clustering algorithm which uses the tolerance rough set based approach to find the different meanings of the query in web search results and then organizes these results into different clusters according to their related meanings about query. Each meaning of the query can be represented by its contexts in each result and if there is a significant correlation between two context words, it is more likely that these two words represent the same meaning of query and also suitable as good indication of the meaning of query. In this study, the search results are organized in groups that each group of results relates to context words with high correlations and then these groups are merged into the final clusters representation using both cluster contents similarity and cluster documents overlap. The correlated context words with high documents coverage are selected as the labels of each cluster. Some experiments were conducted on different search results sets based on various queries. The results and comparisons of the proposed algorithm with that of the popular search results clustering algorithms through an empirical evaluation establish the viability of this proposed approach.
A Study on Part-of-Speech Conversion in English-Chinese Translation of Petroleum Scientific and Technical Texts  [PDF]
Yule Sun
Open Access Library Journal (OALib Journal) , 2026, DOI: 10.4236/oalib.1115610
Abstract: Part-of-speech conversion is one of the core techniques in the translation of petroleum scientific and technical English. In the context of petroleum technical texts, English and Chinese differ remarkably in lexical structures, making literal word-for-word translation hardly feasible. Combining practical translation cases in the petroleum technical corpora (covering drilling, reservoir engineering, oil production, industry standards and field operation manuals), this paper systematically analyzes the conversion rules of nouns, adjectives, prepositions, verbs and other parts of speech in English-Chinese translation. It proposes a three-tier progressive processing strategy consisting of literal retention, part-of-speech conversion and sentence restructuring, and summarizes the systematic application of part-of-speech conversion methods in petroleum scientific and technical translation.
第1页/共1000条
每页显示 条


Home
Copyright © 2008-2020 Open Access Library. All rights reserved.