THE THING ABOUT IT IS.

How We Know.

We don't just report what people say. We test it.

The short version

When someone says an event proves their opponents are lying, evil, or out to get them, we don't just repeat it and move on. We write the claim down exactly as they made it, figure out what else would have to be true if it were right, and go check. Then we tell you what we found, even when it's messy.

The claim

This is the "logic check" you see on every card.

State the claim

We write down exactly what's being said, in the plainest version and in the widest one the same words reach. The plain version has to be fair: the argument as the person making it would recognise it, not a caricature we can knock over.

Write it down twice

These are two settings on the same dial: how much the claim asks you to believe. "Executives are cutting jobs to get richer" can mean they expect to profit from automation, which is modest and checkable. Read as far as it goes, it means profit is the whole point and everything else they say is a lie. We write both down before checking anything, and every check says which one it's testing.

Widest is not the same as strongest. The strongest version of an argument is its most defensible one, and the widest version of a claim is usually its least. We're putting the claim at its biggest, not at its best, and the two pull in opposite directions. Failing a widest-reading check is a much weaker result than failing a plain one, and you should be able to tell them apart at a glance.

This turned out to matter more than we expected. Across the thirteen AI arguments we've published so far, 26 checks tested plain readings and not one came back false. 24 checks tested widest readings and not one held up. People are mostly right about what happened and mostly wrong about what it proves.

Read that gap carefully, including against us. We decide where the widest reading sits, and drawing it further out makes it easier to fail. It isn't evidence that people are wrong; it marks where a claim stops being the kind of thing anyone can check. That boundary is the useful part, and it's the part most arguments never locate.

Ask what else would have to be true

If the claim is right, some other things should be checkable in the real world. We list those things before we go look, so we can't quietly move the goalposts once we know the answer.

Go check

Each thing gets checked against what's actually observable, with a source attached. Sometimes it holds up. Sometimes it doesn't. Sometimes it's a mixed bag.

Give the honest verdict

Holds up, mixed evidence, falls apart, or there's nothing here we can actually test. We show our work on every card either way, not just the checks that make one side look bad.

Two different jobs

Not everything here is a claim you can test, so not everything gets checked the same way.

Tested claims get a grade

Someone argues that an event proves something about the other side. There's a claim, so there are checks, so there's a letter. That's the Fault Line Report.

Recorded statements get a severity

Someone said a thing about a group of people. Whether they said it is settled, so there's nothing to check. Where it lands on the ladder below is our call, and that part is arguable. What matters is how far it goes, so those carry a severity instead. Four steps. Stigmatizing casts a group as a problem to be managed. Hateful treats a whole group as lesser, dangerous, or not fully human. Threatening warns a group of harm, demands they stay quiet, or tells someone to act against them. Calls for violence asks for harm to come to a group, or celebrates harm that's already been done to them. Those cards carry no grade and no illustration, in the Weather Report and in the investigations alike.

Where the sections split

The Fault Line Report is camp against camp: claims people make about each other's motives, which can be tested. The Weather Report is a weekly pass over the same tracked accounts for language aimed at a group. Investigations are longer passes over one question, and each one opens onto every item it found. The World Cup and religion investigations are scored as hate rather than as polarization, so they carry severities, not grades.

Two tracks, kept apart

Hate aimed at a protected group and dehumanizing language aimed at political opponents are both counted, and neither total is ever presented as a score against the other. They're different things with different histories and different consequences. One place they do share a denominator is the severity chart, which asks how far a statement goes rather than who it was aimed at; anywhere the question is who, the tracks stay separate.

We don't reprint slurs

Where the wording is itself a slur, we describe it and link to the source rather than reproducing it. The source is one click away, so you can read it and decide for yourself whether we characterized it fairly. Documenting hate doesn't require handing it a bigger audience than it had.

The result

Two results per card, and they never get averaged into one.

Each reading is reported on its own

Every card shows how the claim fared read plainly and how it fared read at its widest, with the count behind each. Holds up means everything settled came back true. Fails means none of it did. Mixed is everything between.

Why there is no single score

A plain-reading check asks whether the claim describes something real. A widest-reading check asks whether its largest possible version is true. Those are different questions and they don't share a scale, so combining them produces a number whose value depends on how many checks of each kind we happened to write rather than on the evidence. Two cards here have a flawless plain reading and would have landed several letters apart for that reason alone. So the two lines stay separate, and a reader can see which one moved.

Some checks can't be settled, and those leave the count entirely

Can't be checked

Sometimes the world offers no way to answer a question in either direction. Nobody has measured where the gains from AI actually landed. Nobody outside one man's head knows which voices moved him. A strategy meant to stay quiet wouldn't produce a confession even if it were real. Marking those as failed checks would be claiming we looked and came up empty, which is a much stronger statement than the truth. So they drop out of both lines, and the card says which ones and why.

What this isn't

It isn't a truth score, and a claim that doesn't hold isn't a claim that someone lied. It measures how a claim fared against the specific checks we chose to run. We pick those checks, and a different reader could reasonably pick different ones and land somewhere else. That is why every check is printed with its own reasoning and its own source, instead of a number you would have to take on trust.

The scoring

Once a story clears our bar, it gets scored 1 to 5 on five things. You'll see these as chips (like "R 4 · I 4 · P 3") on the full-scoring section of every card.

R · Reach

How far it actually traveled. A senator's floor speech reaches further than a reply nobody saw.

I · Intensity

How nasty the framing gets. Arguing the merits is a 1. Assuming the other side is lying by default is a 3. Treating the whole camp as one evil essence is a 5.

P · Penetration

How many separate, unconnected people actually believe it. One loud account scores low. A belief that's become the camp's default answer scores high.

E · Elite carriage

Whether people with real power (senators, cabinet officials, network anchors) say it themselves, or whether it stays on anonymous accounts.

L · Lock-in

Whether the other side has built a mirror-image version of the same story, so both sides now point at each other's worst moment as proof they were right all along.

Those five combine into a "charge stage," also shown as a chip:

C1 · Something happened

An event occurs. Coverage is still just describing it. Nobody's claimed it as proof of anything yet.

C2 · Both sides start spinning it

Competing "here's what this really shows about them" versions show up. The event becomes ammunition.

C3 · One version wins

One framing becomes the camp's go-to answer. Saying otherwise inside the camp starts to cost you something.

C4 · It outlives the event

The story sticks around and gets reused on whatever happens next, whether or not it actually fits. This is most of what you'll see on this site, because it's the stuff that keeps showing up.

Double-checked

Every quote gets traced to a dated, linkable source rather than a screenshot or someone else's summary. Then a second pass tries to break the finding on purpose: wrong date, missing context, a quote that doesn't actually say what it's being used to say. What survives both passes is what makes it onto this site.

Two kinds of receipt, and we label which is which

Some quotes we read where the person published them: their own post, their own press release, their own filing. Others reached us through a news outlet that was in the room, and those say so directly in the byline, like "as reported by Fortune." We think both are usable, but they aren't the same strength of evidence, so we never blur them.

A verified quote isn't a verified claim

Confirming that someone said a thing, on a date, is a different job from confirming the thing is true. The receipts section does the first. The logic check does the second. A card can have airtight quotes and still grade badly, and that's the instrument working, not failing.

Want the whole thing

This site is the plain-language cut. Every page carries a link to the research it came from, and those pages hold the full carrier sets, the per-item scoring, and the coverage notes on what we couldn't verify.