A collection is a named slice of AltF Lexicon defined by a rule rather than by a hand-typed word list. 199 of them cover the 147,478 entries in the corpus. Most membership comes straight from WordNet’s own lexicographer files and domain labels — the classification a lexicographer applied to each sense. The rest comes from properties of the word you can check yourself: how many letters it has, how many syllables, where the stress falls, how many meanings it carries, how often it appears in everyday English.
Nothing on this page was typed out by hand. Every collection is a predicate run over the whole corpus, in three descending orders of authority, and every collection page repeats its own rule at the top so you can judge the list before you trust it.
Every WordNet sense is filed by a lexicographer under one of 45 subject files — noun.animal, verb.motion, noun.food. If a sense is filed under animals, the word is an animal. That is a classification decision made by a person reading the sense, not a keyword match against the definition text.
Senses also carry topic, region and usage pointers — the field a term belongs to, where it is spoken, how formal it is. Every label with at least 25 member entries becomes its own collection, which is where "Indian English", "slang" and "law vocabulary" come from. 108 of the 199 collections here are generated this way.
Length, repeated letters, missing vowels, syllable count, stress position, sense count, frequency band. These are objective and checkable: you can count the letters in a palindrome and see whether we were right.
Each list is ranked by how common its words are, so the vocabulary you half know sits at the top and the specialist tail sits below it. Lists store up to 600 words on the page; where a collection is larger than that, the page says so and gives the true total.
What the word is about, taken from WordNet's own classification of every sense.
How and where a word is used — the field, the region, the level of formality.
Curiosities of spelling: length, repeated letters, missing vowels, symmetry.
Syllables, stress and the words that are harder to say than to read.
Slices built for people studying English rather than browsing it.
Subject classification, domain labels and semantic relations come from WordNet. Syllables and stress come from the CMU Pronouncing Dictionary where a word is recorded in it. Commonness comes from a frequency list built on everyday English. Full sources and licences.