See how words choose their neighbors.
PowerConc 2 analyzes how words co-select other words and texts. It offers the core functions of WordSmith Tools and AntConc in a simpler interface for users new to corpus linguistics and text analytics.
Explore the tools| a general-purpose | corpus | analysis tool that |
| load a reference | corpus | to compare terms |
| each file in the current | corpus | is checked and removed |
| mixed-script documents in one | corpus | use separate rules |
| underuse in the target | corpus | shows as a negative score |
Start with your own texts
Import plain text files, set up subcorpora, then move into word lists and concordances.
Open Files imports TXT files. Open Folder imports text files from a folder and all its subfolders.
Click Set Subcorpora to load a two-column, tab-separated mapping of filenames to subcorpora. File extensions are optional. Remove deletes all checked files from the current corpus.
Supported encodings:
Auto detection tries UTF-8 first, then falls back to Windows ANSI.
Multiple tools, one workflow
PowerConc 2 is built around two main tools: List and Search. Keyword extends List, and Collocation extends Search. Click a row in List or Keyword results to jump to a KWIC search for that item.
List
Build N-gram lists from your corpus, then compare them against a reference corpus.
N-Gram
Generates N-grams of one to five tokens. Each entry shows frequency, frequency per million N-grams, file count and subcorpus count.
N-grams stay within a single file but may cross sentence boundaries. Distribution is a central part of the design, inherited from PowerConc Version 1, developed by Yunlong Jia in Delphi.
Keyword
Click List to reveal Keyword. Load a reference corpus and compare terms with signed G², χ² and Log Ratio. Negative scores mean a term is underused in the target corpus.
The By File view also reports file-level G² statistics.
Search
Look up words and phrases in context, then examine what they habitually co-occur with.
KWIC
Generates KWIC concordances. Enter a query, choose Regex or literal matching, set case sensitivity and context span. Double-click a result to see an expanded context with the match highlighted.
Terms match complete tokens. N-grams match token sequences, including those that span line breaks.
Collocation
Click Search to reveal Collocation. Set left and right context windows, then read LL, Delta P, MI, MI3, T-score and Z-score. Click a row to see that token in the concordance lines.
See Collocation Statistics below for how each measure is computed.
Token Definition
Custom Regex is the default mode. White Space is available as an alternative.
Chinese texts
[\u4e00-\u9fa5A-ZA-Za-za-z0-90-9\.%%]+
English and other languages
[A-Za-z0-9-]+
For other languages, edit the pattern, for example \p{L}+ for Unicode letters. A tokenization regex must not match empty strings.
Any character in the U+4E00–U+9FA5 range triggers the Chinese rule, even in mixed-script documents. Detection runs once on the full decoded text at import, and every tool then applies the rule per document.
When Chinese texts are loaded, the tokenization controls are locked. Documents not identified as Chinese use the English default. With no Chinese texts, the pattern stays editable, and opening new files restores automatic selection.
Regex engine: PCRE2 with Unicode support. Examples:
\p{Han}+
\bcorpus\b
get(?:s|ting)?
Invalid patterns show an error dialog in English. Match and recursion limits guard against expensive expressions. TSV files are treated as plain text; tagged TSV and POS views are no longer supported.
Collocation Statistics
Six measures per collocate, presented in this order: LL, Delta P, MI, MI3, T-score and Z-score. Results sort by descending Log Likelihood.
| Collocate | LL | Delta P | MI | MI3 | T-score | Z-score |
|---|---|---|---|---|---|---|
| linguistics | ↓ sorted | — | — | — | — | — |
Layout shown for illustration. No sample data is included.
Log Likelihood and Delta P come from a 2 × 2 table of unique corpus tokens: collocate versus other token, inside versus outside the union of node windows.
- Log Likelihood is reported as unsigned G².
- Delta P shows N/A when every token falls inside a node window.
- MI, MI3, T-score, Z-score and positional counts keep repeated occurrences, so overlapping windows are counted repeatedly.
Vocabulary Levels
Set List Mode to Vocabulary Levels to analyze vocabulary lists in TXT, TSV, CSV or XLSX format. Use one file per level, with no header row. The first column holds the family name; later columns hold its members.
A word found in several levels goes to the first level where it appears. Matching ignores case.
Filter, Sort and Export Your Results
- Click a column header to sort.
- Apply regex filters to the current result set, one after another.
- Exclude reverses a filter so matching rows drop out.
- Reset Results restores the full set. Resample randomizes the displayed rows for a specified amount of rows.
- Save writes all filtered results to a UTF-8 CSV file with BOM.
- Chart shows the top 20 values from a numeric column and exports a PNG image.