When New LLMs Impress, Their Flaws Quickly Disappoint
The author tests CodeX and GPT5.6‑Sol on a multi‑task directory workflow and finds simple yet puzzling errors, then observes Fable5 failing on basic CSS tweaks, linking both issues to catastrophic forgetting and hallucination in large language models.
