Same LRU Cache Code, GPT‑5.6 Sol, Terra, and Luna Produce wildly different results
Running an identical LRU‑cache implementation request through GPT‑5.6's three models shows Sol generating 592 tokens with full tests and thread safety, Terra 315 tokens with basic functionality, and Luna only 115 tokens lacking docs and tests, leading to a cost‑benefit analysis that favors Sol for production code despite its higher token price.
