mechanism

REPORTS

Every model release, run through the same pipeline.

Real agents, mostly open source. Every eval case runs five times on each model version, the checks are deterministic rather than a judge model, and a repair only counts once the whole suite has been rerun. Sometimes the answer is do not upgrade.

  1. September 8, 2026 · OpenAI

    GPT-5.5 to GPT-6 Astra

    gpt-5.5gpt-6-astra

    • shell_gpt (upstream a082bd5)SAFE WITH PATCH

    GPT-6 Astra rejected every tool-calling request shell_gpt made. One endpoint change brought the whole suite back.

  2. September 1, 2026 · Anthropic

    Claude Fable 5 to Claude Fable 5.1

    claude-fable-5claude-fable-5-1

    • Cookbook SMS botSAFE WITH PATCH
    • Quickstarts agentNO REGRESSION DETECTED
    • FACTBROKEN BEFORE AND AFTER
    • claudette toolloopNOT RUN LIVE

    Four open-source Claude agents through the Fable 5.1 migration on release day: one was fully repaired, one showed no regression at all, one was already broken on Fable 5, and one was never run live.

  3. August 29, 2026 · OpenAI

    GPT-5.5 to GPT-5.6-sol

    gpt-5.5gpt-5.6-sol

    • shell_gptSAFE WITH PATCH
    • Booking agent (internal stress run)STAY PINNED

    GPT-5.6-sol rejected every tool call shell_gpt made, and one endpoint-routing line brought all 14 cases back. The same pipeline on our own 38-case booking agent could not repair four regressions, and returned STAY PINNED.

Next release: same pipeline, another page.

MIGRATION SESSION

Choose a time.

15 minutes on a migration that broke your agent. Bring the working version, the target version, and whatever evals you have.