Methodology

Every module on a report page carries one of three provenance badges: SSA DATA (computed from public records), AI-WRITTEN (editorial prose, machine-drafted and lint-gated), or VOTES (aggregated live human ballots). This page explains what sits behind each badge.

Data sources

Birth counts come from the U.S. Social Security Administration's public baby-names dataset (ssa.gov/oact/babynames), covering 1880 to the latest published year. Names with fewer than five births in a given year are omitted at the source, and the SSA notes the raw files are unedited — we filter documented administrative placeholders (such as “Unknown”) before publication. Survival weighting uses the SSA actuarial period life table (ssa.gov/oact/STATS). Surname frequencies in the paid rarity estimate come from the U.S. Census Bureau's 2010 surnames file (census.gov).

Estimated age, and why it says “estimated”

The age profile weights each birth-year cohort by its probability of being alive today, taken from the life table. The result is an estimate about the population of living bearers — a distribution, never a fact about any individual. The same survival weighting produces the generation split. Peak years and ranks are read directly from the records, not estimated.

Rarity estimates (paid reports)

Full-name rarity multiplies the estimated living count of a first name by the frequency of a surname. We publish ranges and bands, never a single number — the inputs are estimates, and pretending otherwise would be dishonest. Surnames outside the Census table get a band only.

The Opinion Ledger (votes)

Votes come from visitors, one tap at a time, with no account. We rate-limit by device and network, apply cooldowns, and silently discard ballots that look automated. Crowd results display only after a minimum sample (an early reading appears from 10 ballots, labeled with its sample size; full display at 30). Voters are self-selected internet visitors — the ledger measures perception among people who vote on name websites, and claims nothing more.

AI-written sections

Prose modules marked AI-WRITTEN are drafted by a language model from the statistics above, then machine-checked against an editorial rulebook: banned-cliché lists, a prohibition on claims about any individual person, no invented statistics, and no assertions about intelligence, appearance, morality, ethnicity, religion, or class. Drafts that fail are regenerated or discarded. Factual claims outside our datasets (name origins, notable bearers) are generated conservatively — instructed to omit anything uncertain — but can still contain errors; corrections via the support address are welcome. Content is versioned, and prompt or model changes trigger regeneration and a fresh human sample review.

What we do not do

No astrology, no numerology, no personality claims about you. A report describes a name's statistical footprint and its public perception. People outrank their paperwork.

Questions about the method: see About for who runs this, or write to the support address listed there.

Support: [email protected] · Pricing