MiniCONE现代汉语平衡语料库简介

MiniCONE现代汉语平衡语料库（简称“MiniCONE语料库”）是CONE现代汉语平衡语料库（简称“CONE语料库”）的精编版，总规模约300万词。CONE语料库总规模约5,000万词，相关介绍见：https://corpus.bfsu.edu.cn/CONE.html。

MiniCONE语料库总库容为3,071,626词，其中MiniOralCONE为1,002,993词，MiniNetworkCONE为1,007,142词，MiniEditedCONE为1,061,491词。三个MiniCONE子语料库中的语料文本产生于2018—2022年间。每个子库中分别包含未分词（RAW）、分词（TOK）、词性标注（POS）等三个版本。分词所用工具为HanLP（FINE_ELECTRA_SMALL_ZH），词性标注集为PKU（参见PKU_POS_Tagset.pdf或https://hanlp.hankcs.com/docs/annotations/pos/pku.html）。
MiniCONE语料库总库容根据 HanLP分词标注程序未标记为标点符号的字符串总数计算得出（即词性代码不为 "w" 的字符串）。

MiniCONE语料库延续了CONE语料库“口头汉语（Oral）—网络汉语（Network）—书面汉语（Edited）”三位一体的建库理念，由三个规模相当的子库构成，每个子库约100万词。三个子库在语料采集与抽样设计方面分别参考了国际英语语料库（International Corpus of English, ICE；Greenbaum & Nelson, 1996）的口语采样框架、在线英语语域语料库（Corpus of Online Registers of English, CORE；Biber & Egbert, 2018）的网络语言采样框架以及布朗语料库（Brown Corpus；Francis & Kučera, 1964）的书面语采样框架。

各子库的具体语料来源及构成信息详见语料库根目录下对应的三个元数据文件（metadata）。有关各子库的规模统计信息，请参阅语料库根目录下的 MiniCONE_Word_Count.xlsx 文件。

语料库引用方式
使用CONE语料库开展研究时，请引用以下文献：
Xu, Jiajin & Mingchen Sun (forthcoming). A Frequency Dictionary of Mandarin Chinese: Core Vocabulary for Learners (2nd Edition). Routledge.

语料库访问地址
http://114.251.154.212/cqp/
用户名：test；密码：test

参考文献
Biber, D., & Egbert, J. (2018). Register variation online. Cambridge University Press.
Francis, W. N., & Kučera, H. (1964). Manual of information to accompany a standard corpus of present-day edited American English, for use with digital computers. Brown University.
Greenbaum, S., & Nelson, G. (1996). The International Corpus of English (ICE) Project. World Englishes, 15(1), 3–15.

*****

Introduction to the MiniCONE Balanced Corpus of Modern Chinese

The MiniCONE Balanced Corpus of Modern Chinese (henceforth the “MiniCONE Corpus”) is a small subset of the CONE Corpus: A Corpus of Oral, Network, and Edited Chinese (hereafter referred to as the “CONE Corpus”), totaling approximately three million tokenized Chinese words. The CONE Corpus contains approximately 50 million words. Further information is available at: https://corpus.bfsu.edu.cn/CONE.html.

The total size of the MiniCONE corpus is 3,071,626 words, including 1,002,993 words in MiniOralCONE, 1,007,142 words in MiniNetworkCONE, and 1,061,491 words in MiniEditedCONE. The texts in the three MiniCONE sub-corpora were produced between 2018 and 2022.

Each sub-corpus contains three versions of the texts: untokenized raw texts (RAW), tokenized texts (TOK), and part-of-speech tagged texts (POS) respectively. The tool used for word tokenization is HanLP (FINE_ELECTRA_SMALL_ZH), and the part-of-speech tagset is PKU (see PKU_POS_Tagset.pdf or https://hanlp.hankcs.com/docs/annotations/pos/pku.html). 
The word token count of the MiniCONE Corpus was calculated based on the total number of tokens that were not tagged as punctuation marks by HanLP (i.e., tokens with a code other than "w").

The MiniCONE Corpus inherits the tripartite design principle of the CONE Corpus, namely “Oral Chinese (Oral) — Network Chinese (Network) — Written Chinese (Edited)”. It consists of three sub-corpora of approximately equal size, each containing around one million words. In terms of corpus sampling design, the three sub-corpora respectively draw on the sampling frameworks of the spoken component of the International Corpus of English (ICE; Greenbaum & Nelson, 1996), the online language sampling framework of the Corpus of Online Registers of English (CORE; Biber & Egbert, 2018), and the written language sampling framework of the Brown Corpus (Francis & Kučera, 1964).

Detailed information on the sources and makeup of each sub-corpus can be found in the three metadata spreadsheets in the root directory of the corpus. Statistics on the size of each sub-corpus are also provided in the MiniCONE_Word_Count.xlsx in the corpus root directory.

Citation
When conducting research using the CONE Corpus, please cite the following publication:
Xu, Jiajin & Mingchen Sun (forthcoming). A Frequency Dictionary of Mandarin Chinese: Core Vocabulary for Learners (2nd Edition). Routledge.

Corpus Access
Access URL:
http://114.251.154.212/cqp/
Username: test
Password: test

References
Biber, D., & Egbert, J. (2018). Register variation online. Cambridge University Press.
Francis, W. N., & Kučera, H. (1964). Manual of information to accompany a standard corpus of present-day edited American English, for use with digital computers. Brown University.
Greenbaum, S., & Nelson, G. (1996). The International Corpus of English (ICE) Project. World Englishes, 15(1), 3–15.