When New AI Models Impress, Their Flaws Quickly Disappoint
The author tests CodeX, GPT5.6‑Sol and Fable5, exposing simple yet puzzling errors—missed tasks, inconsistent CSS changes, and contradictory answers—that may stem from catastrophic forgetting and highlight emerging reliability bottlenecks in large language models.
