Skip to content

Compare two setups

A setup is one way of producing labels: a human questionnaire version, or one model under one prompt recipe. Choose two and see how many of the same 3,000 tweets end up labeled differently.

Not sure where to start? Try one of these:

What each setup produces

What changes between them

Which tweets the two setups disagree on

Save or share this comparison

More on the statistics

Tweets that get a different label