whatupwolf · lab · experiment

Reasoning-trace A/B

A confident answer is easy. A trustworthy one is harder to tell apart than you'd think.

An AI gives you a question, a confident final answer, and the reasoning trace it says it used to get there. Your one job each round: decide whether you trust the reasoning.

Some traces are sound. Others reach the same confident answer through a subtly broken step — a fabricated fact, an invented premise, a bad justification that happens to land right. The catch: a wrong reason for a right answer is exactly what's hard to notice. That's the thing this little study is about.

Trace shows the moment the answer does.

  • 4 rounds — a mix of sound traces and subtly unfaithful ones, shuffled fresh each time.
  • Judge each: trust it, or flag it. Flag one and you can point at the step you suspect.
  • At the end: how many flaws you caught, how many fooled you, and what the research found.

Runs entirely in this one file — offline, no accounts, no API key, no live model call. Every trace below is hand-written and fixed.

Built as a lab experiment for whatupwolf.com/lab. One HTML file, no build step, no dependencies, no network calls — the traces are static and hand-written, so it works offline and reads the same on a plane. The finding it demonstrates is real; the example traces are illustrations of it.