Measuring what people don't say.
An implicit association test measures how quickly you sort things into categories when those categories share a key. If two ideas are already linked in your head, pairing them is easy and you are fast. If they aren't, you hesitate — by somewhere between forty and two hundred milliseconds. That gap is the measurement.
This runs the real thing: the seven-block design from Greenwald, Nosek and Banaji (2003), scored with their improved D algorithm, with the data-quality exclusions applied, a bootstrap interval around the estimate, and a written summary that is never allowed to see a raw latency — from a language model when one is configured, and from a deterministic rules engine when it is not.
1 · Choose a study
Advanced — custom brands and a fixed seed
The seed decides the trial order and both counterbalancing factors. Two sessions with the same seed present exactly the same stimuli in exactly the same order — which is what makes a session re-analysable rather than merely repeatable.
2 · Before you start
You will see brand names and descriptive words one at a time. Press E for the category shown on the left and I for the one on the right. Go as fast as you can while staying accurate — a wrong key shows a red ✗ and you must correct it before the next trial.
Use a physical keyboard, sit somewhere without interruptions, and keep this tab in the foreground. The task detects and reports it if you switch away.
The worked example runs a synthetic respondent through the same pipeline so you can see the full report without doing 190 trials. It is labelled as simulated everywhere it appears.
What this deployment is doing
v—Scoring the session
This takes a few seconds — the interval is bootstrapped from your own trials.
—
Sensitivity to the scoring choice
The same trials scored under other defensible published procedures. If the conclusion moves when the algorithm changes, that is a fact about the data, not a footnote — so it is reported next to the headline rather than buried.
Where the time went
Correct-trial latencies, 50 ms bins shared between the two conditions so the bars are directly comparable. Reaction times are right-skewed — that long tail is why the analysis uses resampling rather than a t-test.
Trial by trial
Every scored trial in the order it was presented. Practice effects, fatigue and the block boundaries are visible here and invisible in a mean.
Is this distinguishable from chance?
—Block-level detail
Executive summary
— —View the exact prompt and the data the model was given
If a model writes part of a research report, the instructions it was given are part of the method and belong in the audit trail.