Context runs entirely in your browser
0 tokens 0 types 0 texts

What can words reveal?

Explore TV characters, investigate Arabic judicial language, or bring your own texts. Find verbal fingerprints, recurring phrases, concept maps, patterns, and hidden themes.

Visual Lab

See semantic neighbourhoods, phrase flows, sentence rhythm, and corpus contrast from several angles.

Everything is computed here. The projection, graphs, tagging, and layouts run locally in this browser; hover over marks to inspect them.

Semantic map

t-SNE projection of word-context PPMI vectors

Multi-hop collocation flow

Focus → strongest neighbours → their strongest neighbours

Sentence rhythm

Hover to inspect · click to pin a sentence
Move across the waveform to inspect sentence length and text.

Contrast cloud

Cool = over-used here · warm = over-used in the reference

Explore change and contrast

Follow words across ordered texts, compare their intensity, and—when dialogue is detected—investigate who speaks, how, and with whom.

Character Lab. Add scripts formatted as Speaker: dialogue to unlock speaker comparison automatically.

Word frequency

Raw counts alongside a normalised rate per million words, which is what lets you compare texts of different lengths.

Frequency list

N-grams and clusters

Recurring word sequences. Function words are kept by default — they are what make a cluster a cluster.

Multi-word expressions

Beyond raw frequency. Candidates are ranked by association strength and document range, separating cohesive expressions from sequences that merely happen to be frequent.
Phrase patternsExplore variable-slot phrases and compare repeated sequences with a reference corpus.

Keyword in context

Every occurrence of the node with its surrounding text. Sorting by the first word to the left or right is how patterns become visible.

Regex / CQL query builder

Search syntax guide ↗
Pattern and batch searchCombine CQL, regular expressions, grammatical presets, filtering, and multiple queries.
Syntax guide ↗

Collocations

Words that occur near the node more often than chance would predict. The measures disagree by design: MI favours rare, tightly-bound partners; t-score and log-likelihood favour frequent, reliable ones.

Reading the table. Obs is how often the collocate falls inside the span; Exp is how often it would if words were sprinkled at random. L/R shows positional preference — a collocate that only ever appears to the left is usually a modifier, one that only appears to the right is often a complement.

Collocation network

The node's strongest collocates, plus the links among the collocates themselves. Clusters here often correspond to distinct senses or phraseologies of the node.

Association ExplorerExpand several node words, reveal shared collocates, and record the active settings.

Keyness

What this corpus over- and under-uses relative to a reference corpus, by log-likelihood.

No reference corpus loaded

Keyness is a comparison, so it needs a second corpus to compare against. Add reference texts in the left panel and analyse again.

Side-by-side corpus difference

Corpus comparisonInspect stable words, exclusive terms, effect sizes, and overlap alongside keyness.

Dispersion

Whether a word is spread evenly through the corpus or concentrated in a few places. A high raw frequency driven by one burst tells you something different from the same frequency spread throughout.

Dispersion of the most frequent words

Measures. Range counts the parts a word appears in at all. Juilland's D and Carroll's D₂ run from 0 (all in one place) to 1 (perfectly even). DP is the reverse — 0 is even, 1 is maximally clumped.
Distribution profileCompare occurrences and normalized rates across every text, with a compact statistical summary.

Parts of speech

English uses an offline rule-based tagger. Other supported languages use transparent lexical function-word classes.

Distribution

Most frequent word by class

Accuracy caveat. These offline classes are designed for exploration, not publication-ready annotation. Treat the counts as indicative and spot-check anything you plan to report.

Readability & sentence structure

Six standard formulas, plus the sentence-length distribution they are all ultimately derived from.

Sentence length distribution

How these are computed. Every formula here depends on counting syllables, which is done with a spelling heuristic rather than a pronunciation dictionary. It handles the large majority of English words but undercounts hiatus (create, poem). Grade-level scores are best read as approximate, and as comparisons between texts rather than absolute claims.

Word cloud

Size is proportional to frequency. Useful as a first impression; the frequency table is what you should quote.

Principal component analysis

Each point is one analysis unit, positioned by its tf-idf profile. Units that sit near each other use similar vocabulary.

Term loadings

Terms pulling hardest on each component.
Implementation. Power iteration with deflation on the centred document–term matrix, not a full SVD. The leading components are accurate; later ones accumulate error. Explained-variance figures are relative to the total variance of the term matrix.

Topic modeling

Non-negative matrix factorisation over the tf-idf matrix. Each topic is a weighted set of terms; each unit is a mixture of topics.

Topic composition by unit

Topic constellation

Bubble size = corpus share · proximity = vocabulary similarity

Evidence in context

Select a topic card or bubble
When not to trust this. Topic models need many reasonably long, genuinely distinct documents. On a single short text split into segments, "topics" will often be artefacts of where the cuts fell rather than real thematic structure. Check the segment count reported above before reading anything into the result.

Analysing…