How we count
The rules behind every figure.
A researcher should be able to defend any number in a meeting. These are the rules the desk follows, written down so you can. They are the same in the free sample, in your browser and on the desk.
Binding a question to a field
When a study is added, every question in the file is read once. The desk notes its wording, the words it also answers to, its grain (a parent question or one item of it), its tense (past, present, hypothetical) and its polarity (asked the positive way or the negative way). Those notes score every English question you type afterwards.
A question you ask is matched against every field and the fields are ranked. The top field is used only when its score clears a floor and it leads the runner-up by a clear margin. A question about last season is scored against the tense of each field, so it does not land on next season’s look-alike.
A pin is your decision, remembered. When you pin a wording to a field, every later ask with that wording goes to that field, and the card says it was pinned.
When there is no figure
- The question names something the study never asked about. The card says what was not found, and gives no figure.
- No field clears the floor. A nearest column is not used.
- Two fields score too close together. The card shows both and asks you to pin one. Your client sees “too close to call” and is told to ask you.
- The field exists in the study but has no column on this wave. The card stays blank on that wave and names the look-alike that is in the file, with the percentage it would have given.
- A narrowed base falls under five people.
A refusal is a result. It gets a receipt and it can be sent to the client page as a blank.
The base
A percentage is counted on the people who answered that
question. Blanks are not in the base. The card shows the base as
n = 100 and the sentence says “of 100”.
“Among those who …” narrows the base to the people who gave one answer to another question. The card says who, and how many. A base under five people is refused rather than shown.
Weights and effective bases
A weight column in the file is applied and named. The weighted figure leads and the unweighted figure stands beside it, so nobody has to ask which one they are looking at.
Alongside the weighted figure the desk shows the effective base: the sum of the weights squared, divided by the sum of the squared weights. That is the base the significance tests run on when a weight is in play, so a heavily weighted cell is not treated as if it had more people than it does.
Cuts, column letters and the 95% test
Add “by region” and the figure is counted in each column of that question. Every column gets a letter. A cell lists the letters of the columns it is higher than at the 95% level, the way agency tables have always carried them.
- Shares are tested with a two-proportion z. Means are tested with a Welch z on their spreads.
- Weighted figures are tested on their effective bases when every column has one; otherwise on unweighted bases. The footnote under the table says which.
- A column under five people shows “Too few people” instead of a figure, and no test is run on it.
Scales
A 1–5 agreement question or a 0–10 likelihood question is read as a scale. The desk works out which end is favourable from the answer labels. The card reports the favourable box, top 2 on a five-point scale and top 3 on 0–10, and the mean, both counted on the people who answered on the scale. A “don’t know” code sits outside the scale and is left out of both.
Ask for the mean and the card leads with the mean. Ask for the NPS on a 0–10 question and the card leads with the NPS. A frequency scale (“once a week”, “every day”) has no favourable end and stays a plain count.
Amounts
A weekly spend, an age, a number of children: the one figure that makes sense on an amount is the mean, on the people who gave a number. Codes that stand for “no answer” are left out by name. The desk takes them from the file where it says so (SPSS user-missing, a label like “Refused”), and otherwise recognises the two conventions every agency uses: all nines above every real value, and a negative code on an amount that is otherwise never negative. The card says which codes were excluded.
Who was asked
A questionnaire asks some questions only of some people. The response file does not carry that logic, so most tools show “of 459” and leave the reader to look up why. The desk reads the routing back from the pattern of blanks: for each question with blanks, it looks for the question whose answers explain them. When the rule fits, the card says so in a sentence: “Asked only of the 484 who ticked the brand on the awareness question.” When the fit is good but not exact, it says “mainly” rather than “only”. When no rule fits, nothing is claimed.
The reading band
A figure rests on choices a reader never sees: whether the don’t-knows sit in the base, and which of two look-alike questions was meant. A chat picks silently. A crosstab tool makes you pick. The desk counts the figure the one way the question asked for, then counts it again each other reasonable way, and says whether it held.
“With don’t-knows in the base it is 35.4%. The figure holds.” is a number that can be used as it stands. “Careful: if you meant the unaided question it is 27.8%, against 13.5% as counted.” is a number whose wording should be read out before it is used. Only a look-alike the study itself names counts as another reading. The receipt stays on the one reading the question asked for.
Waves
Each wave is its own file. A question is counted on each wave where it has a column, and nowhere else. With two waves the card marks the move between them: “+4.2 points from Last season to Next season. A real change, at the 95% level, on 100 and 100 people.” or “Within what chance would give.” The test is a two-proportion z for shares and a Welch z for means, on effective bases when both waves carry a weight. With three or more waves the card draws the line.
The mark between waves is a reading aid. It is not part of any receipt.
The receipt
Every result, counted or refused, gets a receipt: a SHA-256 hash
over the study, the wave, the question as typed, the field it
bound to, the base, every answer row with its count and
percentage, a fingerprint of the codebook, and
model=none. Where a cut, a narrowing, a weight, a
scale or an amount was involved, that is in the hash too.
Run the same question on the same file a year later and the
hash matches. Change one row and it does not. Anyone with the
receipt can re-hash it at /verify on the desk,
without a login. The response file is never part of the
receipt and never leaves your login.
Ask across every study on your login and each study keeps its own receipt. There is no blended hash because there is no blended figure.
See the rules at work.
The Harborline sample on the homepage runs this exact engine in your browser. Ask it for a cut, a narrowing, a mean. Or drop a file of your own and nothing is uploaded.