Scaling Model Width Alone Hurts AUC: Two Seniors' 1000-Experiment TAAC 2026 Breakthrough
Two senior undergraduates won the Scaling Law Original Breakthrough Award and 10th place in the academic track at the 2026 Tencent Advertising Algorithm Competition by running nearly 1000 experiments, discovering that increasing model width alone reduces AUC unless accompanied by new informative tokens, leading to their CoSFormer architecture that scales capacity and information jointly.
Background
The 2026 Tencent Advertising Algorithm Competition (TAAC) featured a team named "gg" consisting of two senior undergraduate students participating in a recommender-system competition for the first time. They ultimately won two awards: the Scaling Law Original Breakthrough Award and 10th place in the academic track.
Initial Struggles and Data Preparation
The team started poorly, ranking low for several days without a clear optimization direction. They later gained insights from forum posts by other contestants. A key retrospective lesson was that they had not thoroughly analyzed the demo data provided by the organizers before modifying models, likening it to "starting a journey before calibrating the map."
Rapid Iteration with GPU Pipeline
During the final round, the team used 7 GPUs, with each model version taking 2–3 hours to train. They often did not wait for full convergence; the trend was visible early. This speed created a tight loop: while one version trained, they read papers, reviewed the previous run, and prepared the next experiment. The GPUs never idled, and the team maintained continuous iteration until the competition ended — breaking into the academic top 10 only in the last two days.
Scaling Experiments: Width Alone Fails
In the final days, the team focused on scaling. One striking finding: keeping the original tokens and only increasing model width caused AUC to drop instead of rise. Reproducing the experiment confirmed the result. The analogy: adding more desks to an office without new documents does not increase the work that can be processed.
CoSFormer: Joint Scaling of Capacity and Information
The team did not abandon width scaling. Further experiments showed that expanding model capacity while supplying more conditions and information — and raising dimensionality in sync — continued to improve performance. Their proposed CoSFormer organizes static features, explicit cross features, and sequence-related features into distinct token types. The model then expands alongside these new tokens, so the added width learns fresh, effective information rather than reprocessing the same inputs.
FFN Sharing Strategy: Grouped by Semantic Type
With multiple token types, the team faced a design choice: share a single feed-forward network (FFN) across all tokens, give each token its own FFN, or something in between. They tested both extremes: a shared FFN struggled to capture semantic differences between token types, while independent FFNs introduced too many parameters and led to overfitting (performance dropped). The final solution grouped tokens by semantic type — tokens of the same type share an FFN, different types use separate FFNs — preserving shared constraints while giving each information type dedicated processing space. This pattern of testing both extremes and settling in the middle characterized many of their nearly 1000 experiments.
Post-Competition Controlled Experiments and Idle Width Hypothesis
After the competition, the team ran additional controlled experiments to understand why width alone hurts. They compared three settings: (1) increase width only, (2) add tokens without new effective information, (3) jointly expand effective information and width. Results showed that joint scaling yielded stable gains, while the other two quickly hit a plateau. From this they proposed the "idle width" explanation: when input information does not grow with model capacity, the extra width remains underutilized and fails to translate into real improvement.
Limitations and Continuous Investigation
The team acknowledges that the idle-width hypothesis is only validated on this competition's data and experimental setup. Whether it holds on other datasets or after controlling for parameter count and compute remains to be verified. Their culture of scrutinizing every failed run — "rising-score versions are kept, dropping-score results are re-examined" — meant that even after winning, unresolved questions were not abandoned.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Tencent Advertising Technology
Official hub of Tencent Advertising Technology, sharing the team's latest cutting-edge achievements and advertising technology applications.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
