REPORTS
Every model release, run through the same pipeline.
Real agents, mostly open source. Every eval case runs five times on each model version, the checks are deterministic rather than a judge model, and a repair only counts once the whole suite has been rerun. Sometimes the answer is do not upgrade.
GPT-5.5 to GPT-6 Astra
gpt-5.5gpt-6-astra- shell_gpt (upstream a082bd5)SAFE WITH PATCH
GPT-6 Astra rejected every tool-calling request shell_gpt made. One endpoint change brought the whole suite back.
Claude Fable 5 to Claude Fable 5.1
claude-fable-5claude-fable-5-1- Cookbook SMS botSAFE WITH PATCH
- Quickstarts agentNO REGRESSION DETECTED
- FACTBROKEN BEFORE AND AFTER
- claudette toolloopNOT RUN LIVE
Four open-source Claude agents through the Fable 5.1 migration on release day: one was fully repaired, one showed no regression at all, one was already broken on Fable 5, and one was never run live.
GPT-5.5 to GPT-5.6-sol
gpt-5.5gpt-5.6-sol- shell_gptSAFE WITH PATCH
- Booking agent (internal stress run)STAY PINNED
GPT-5.6-sol rejected every tool call shell_gpt made, and one endpoint-routing line brought all 14 cases back. The same pipeline on our own 38-case booking agent could not repair four regressions, and returned STAY PINNED.
Next release: same pipeline, another page.