Skip to content

Every setup

How wide is the set of results across defensible choices? Every one of the 90 setups labeled the same 3,000 tweets. A setup is one model under one prompt recipe, averaged over three runs, or one human questionnaire version.

What moves the label most?

Each bar is the spread in prevalence you get by changing one thing and holding everything else fixed. Longer means that choice matters more.

Loading…

Humans and models on one axis

Where do the five human versions fall relative to the LLM setups?

Prevalence by model and prompt recipe

Numbers are the percent of tweets given the label. Hover or focus a cell for the three run-level prevalences; click it, or press Enter, to open that setup in the comparison view.

Loading…

Table view