See how words choose their neighbors.

PowerConc 2 analyzes how words co-select other words and texts. It offers the core functions of WordSmith Tools and AntConc in a simpler interface for users new to corpus linguistics and text analytics.

Explore the tools

Start with your own texts

Import plain text files, set up subcorpora, then move into word lists and concordances.

Open Files imports TXT files. Open Folder imports text files from a folder and all its subfolders.

Click Set Subcorpora to load a two-column, tab-separated mapping of filenames to subcorpora. File extensions are optional. Remove deletes all checked files from the current corpus.

Supported encodings:

UTF-8 (BOM optional)GB2312GBKGB18030ANSI

Auto detection tries UTF-8 first, then falls back to Windows ANSI.

Multiple tools, one workflow

PowerConc 2 is built around two main tools: List and Search. Keyword extends List, and Collocation extends Search. Click a row in List or Keyword results to jump to a KWIC search for that item.

List

Build N-gram lists from your corpus, then compare them against a reference corpus.

N-Gram

Generates N-grams of one to five tokens. Each entry shows frequency, frequency per million N-grams, file count and subcorpus count.

N-grams stay within a single file but may cross sentence boundaries. Distribution is a central part of the design, inherited from PowerConc Version 1, developed by Yunlong Jia in Delphi.

Keyword

Click List to reveal Keyword. Load a reference corpus and compare terms with signed G², χ² and Log Ratio. Negative scores mean a term is underused in the target corpus.

The By File view also reports file-level G² statistics.

Settings #1

Token Definition

Custom Regex is the default mode. White Space is available as an alternative.

Chinese texts

[\u4e00-\u9fa5A-ZA-Za-za-z0-90-9\.%%]+

English and other languages

[A-Za-z0-9-]+

For other languages, edit the pattern, for example \p{L}+ for Unicode letters. A tokenization regex must not match empty strings.

Any character in the U+4E00–U+9FA5 range triggers the Chinese rule, even in mixed-script documents. Detection runs once on the full decoded text at import, and every tool then applies the rule per document.

When Chinese texts are loaded, the tokenization controls are locked. Documents not identified as Chinese use the English default. With no Chinese texts, the pattern stays editable, and opening new files restores automatic selection.

Regex engine: PCRE2 with Unicode support. Examples:

\p{Han}+
\bcorpus\b
get(?:s|ting)?

Invalid patterns show an error dialog in English. Match and recursion limits guard against expensive expressions. TSV files are treated as plain text; tagged TSV and POS views are no longer supported.

Settings #2

Collocation Statistics

Six measures per collocate, presented in this order: LL, Delta P, MI, MI3, T-score and Z-score. Results sort by descending Log Likelihood.

CollocateLLDelta PMIMI3T-scoreZ-score
linguistics↓ sorted—————

Layout shown for illustration. No sample data is included.

Log Likelihood and Delta P come from a 2 × 2 table of unique corpus tokens: collocate versus other token, inside versus outside the union of node windows.

Delta P = P(collocate | inside) − P(collocate | outside)
  • Log Likelihood is reported as unsigned G².
  • Delta P shows N/A when every token falls inside a node window.
  • MI, MI3, T-score, Z-score and positional counts keep repeated occurrences, so overlapping windows are counted repeatedly.
Settings #3 · List Mode

Vocabulary Levels

Set List Mode to Vocabulary Levels to analyze vocabulary lists in TXT, TSV, CSV or XLSX format. Use one file per level, with no header row. The first column holds the family name; later columns hold its members.

A word found in several levels goes to the first level where it appears. Matching ignores case.

Level TokensLevel TypesLevel FamilyFamily Distribution
Settings #4 · Output

Filter, Sort and Export Your Results

  • Click a column header to sort.
  • Apply regex filters to the current result set, one after another.
  • Exclude reverses a filter so matching rows drop out.
  • Reset Results restores the full set. Resample randomizes the displayed rows for a specified amount of rows.
  • Save writes all filtered results to a UTF-8 CSV file with BOM.
  • Chart shows the top 20 values from a numeric column and exports a PNG image.