Four sessions replayed from the labs' recorded output, including the local-model caveats those labs are candid about: the setup that fails then passes, a token bucket draining to zero, a judge that over-punishes honesty, and a loop that sometimes correctly calls no tool at all.