Tagged articles

Parameter Sharing

3 articles · Page 1 of 1
Machine Heart
Machine Heart
Sep 8, 2026 · Artificial Intelligence

SMELT: Looped Transformers Outperform Baselines Under Fair Budget Matching

The SMELT framework from Tsinghua and ByteDance Seed fairly compares Looped Transformers against baselines by matching compute, parameters, and KV cache, finding that looping the middle 50% of layers twice with a larger depth-width ratio consistently reduces validation loss across scales, saves 6.8–18% training compute, and yields downstream gains beyond loss reduction.

Budget MatchingLLM ArchitectureLooped Transformer
0 likes · 12 min read
SMELT: Looped Transformers Outperform Baselines Under Fair Budget Matching
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jun 10, 2026 · Artificial Intelligence

Beyond Orchestrating Workflows: How UnityMAS-O Trains LLM-Based Multi‑Agent Systems

UnityMAS‑O introduces a general reinforcement‑learning framework that converts predefined LLM multi‑agent workflows into trainable tasks, enabling credit assignment across roles, supporting parameter‑sharing configurations, and demonstrating significant F1 and test‑pass improvements on QA and code‑generation benchmarks.

LLMMulti-Agent Reinforcement LearningPPO
0 likes · 12 min read
Beyond Orchestrating Workflows: How UnityMAS-O Trains LLM-Based Multi‑Agent Systems
DataFunTalk
DataFunTalk
Feb 21, 2021 · Artificial Intelligence

Intra‑Ensemble in Neural Networks

This paper proposes an intra‑ensemble strategy that trains multiple sub‑networks within a single neural network using random training operations, width‑depth variations, and parameter sharing, achieving diverse models and improved performance comparable to traditional ensembles while adding only marginal parameter overhead.

Architecture SearchModel DiversityParameter Sharing
0 likes · 9 min read
Intra‑Ensemble in Neural Networks