Tagged articles

large language models

1419 articles · Page 2 of 15
Machine Heart
Machine Heart
Jul 28, 2026 · Artificial Intelligence

Can GPT‑5.6 Sol Crack Fermat’s Last Theorem After 33 Hours of Continuous Running?

A researcher let GPT‑5.6 Sol run for about 33 hours trying to find a simpler proof of Fermat’s Last Theorem, but OpenAI’s system halted the session, prompting analysis of the model’s self‑diagnosis, safety mechanisms, possible bugs, and the broader implications of restricting powerful AI for high‑stakes mathematics.

AI safetyFermat's Last TheoremGPT-5.6
0 likes · 5 min read
Can GPT‑5.6 Sol Crack Fermat’s Last Theorem After 33 Hours of Continuous Running?
DataFunSummit
DataFunSummit
Jul 28, 2026 · Artificial Intelligence

Designing Next‑Gen Recommendation and Search with Multi‑Agent AI Architecture

The article reviews a series of technical case studies—including Alibaba Cloud AI Search's Agentic RAG, Baidu's GRAB generative ranking, Huawei Noah's LLM‑enhanced recommendation, and Elasticsearch vector RAG—showing how multi‑agent AI architectures address high‑concurrency, multimodal, and multi‑hop query challenges while delivering measurable performance gains.

AI AgentsAlibaba Cloud AI SearchBaidu GRAB
0 likes · 6 min read
Designing Next‑Gen Recommendation and Search with Multi‑Agent AI Architecture
Data Party THU
Data Party THU
Jul 28, 2026 · Artificial Intelligence

Can AI with Pre‑1905 Knowledge Become the Next Einstein? DeepMind Examines the Missing Step

The article reviews DeepMind’s “LLMs can’t jump” paper, arguing that even if large language models are fed all scientific knowledge up to 1905, they still cannot recreate Einstein’s breakthrough because they lack the abductive jump from physical intuition to new axioms, a capability that requires interactive world models and physical priors.

Abductive ReasoningArtificial IntelligencePhysics
0 likes · 8 min read
Can AI with Pre‑1905 Knowledge Become the Next Einstein? DeepMind Examines the Missing Step
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 27, 2026 · Artificial Intelligence

Scaling Residual Streams Efficiently: From DeepSeek mHC to xHC’s 16‑Stream Expansion

The blog details how xHC expands language‑model residual streams to 16, achieving nearly double the gain of DeepSeek mHC on 18B and 28B MoE models, and explains the design of Temporal Feature Augmentation and Sparse Write that make large‑N scaling both effective and affordable.

Hyper-ConnectionsMixture of Expertslarge language models
0 likes · 24 min read
Scaling Residual Streams Efficiently: From DeepSeek mHC to xHC’s 16‑Stream Expansion
ThinkingAgent
ThinkingAgent
Jul 27, 2026 · Artificial Intelligence

The Awakening of Large Models: From Classic Language Modeling to Generative AI

This article traces the 56‑year evolution of language models—from ELIZA’s rule‑based scripts and N‑gram statistics to neural embeddings, RNNs, Transformers and the seven‑layer ChatGPT architecture—explaining why the simple next‑token probability definition has remained the core of generative AI, how autoregressive factorization drives training, generation and decoding, why hallucinations arise, and what engineering trade‑offs matter in production.

ChatGPTLLMautoregressive modeling
0 likes · 26 min read
The Awakening of Large Models: From Classic Language Modeling to Generative AI
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 26, 2026 · Artificial Intelligence

From Visual Compression to Memory: MonkeyOCRv2 Reconstructs Document Evidence

MonkeyOCRv2 shows that a reconstruction‑based visual encoder dramatically improves document understanding by preserving fine‑grained page evidence, as demonstrated through controlled encoder swaps, shuffled‑text tests, CHAOS‑Bench conflict evaluations, and consistent gains across seven downstream OCR tasks.

BenchmarkingDocument AIOCR
0 likes · 17 min read
From Visual Compression to Memory: MonkeyOCRv2 Reconstructs Document Evidence
Machine Heart
Machine Heart
Jul 26, 2026 · Industry Insights

Why Are Leading AI Companies Competing for Top Scientists?

The article analyzes the accelerating migration of elite scientists—including a Fields Medalist—to top AI firms like OpenAI, Anthropic, and ByteDance, highlighting how these companies are reshaping research priorities, benchmark practices, and long‑term AI development strategies.

AI researchBenchmarkingByteDance
0 likes · 11 min read
Why Are Leading AI Companies Competing for Top Scientists?
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 25, 2026 · Artificial Intelligence

Why Anthropic Cut 80% of Claude Code System Prompts Overnight

Anthropic discovered that with the release of Claude Opus 5 the previous, heavily‑engineered system prompts became unnecessary, so they removed more than 80% of Claude Code’s prompts without measurable loss, and outlined new concise, context‑driven best practices for LLM prompt engineering.

AI developmentAnthropicClaude
0 likes · 11 min read
Why Anthropic Cut 80% of Claude Code System Prompts Overnight
DataFunSummit
DataFunSummit
Jul 25, 2026 · Artificial Intelligence

The Hidden Flaws of AI‑Driven “Lights‑Off” Software Factories

While AI‑powered coding agents promise a lights‑off software factory where developers never read code, this article reveals the growing maintainability nightmare, benchmark shortcomings, and why current large‑language models still fail to produce good design, urging a return to planning and human oversight.

AI codingSoftware Factoryagentic development
0 likes · 13 min read
The Hidden Flaws of AI‑Driven “Lights‑Off” Software Factories
Ops Development & AI Practice
Ops Development & AI Practice
Jul 25, 2026 · Industry Insights

How to Counter Claude’s Dominance: Porter’s Competitive Strategies for the Large‑Model Market

The article applies Michael Porter’s three generic strategies and value‑chain analysis to show why Claude’s lead in Terminal‑Bench does not guarantee market supremacy, and how cost leadership, differentiation, and niche focus can enable challengers to thrive in the AI large‑model industry.

AI industryClaudeCost Leadership
0 likes · 7 min read
How to Counter Claude’s Dominance: Porter’s Competitive Strategies for the Large‑Model Market
ThinkingAgent
ThinkingAgent
Jul 25, 2026 · Artificial Intelligence

From Next Token to Deployable AI: A Comprehensive Overview of Large Model Technology

This article maps the entire large‑model production chain—from data collection, token prediction, and architecture design through training, alignment, inference, multimodal perception, agentic action, deployment, evaluation, and safety—highlighting key engineering decisions, trade‑offs, and concrete examples.

Agent SafetyInference OptimizationRetrieval-Augmented Generation
0 likes · 47 min read
From Next Token to Deployable AI: A Comprehensive Overview of Large Model Technology
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 24, 2026 · Artificial Intelligence

Why Large-Model RL Training Narrows Over Time? ACL 2026 Paper Reveals Entropy Collapse

The article analyzes why reinforcement learning with verifiable rewards (RLVR) for large models experiences rapid policy‑entropy collapse, breaks the phenomenon down to token‑level entropy changes driven by clipping, advantage, token probability and conditional entropy, and introduces STEER, a token‑wise reweighting scheme that stabilizes entropy and yields consistent performance gains on math and code benchmarks.

RLVRSTEERentropy collapse
0 likes · 14 min read
Why Large-Model RL Training Narrows Over Time? ACL 2026 Paper Reveals Entropy Collapse
DataFunSummit
DataFunSummit
Jul 24, 2026 · Industry Insights

Why High-Quality Data Is the New Bottleneck in Large Model Competition

In a four‑hour investor briefing, DeepSeek founder Liang Wenfeng explains that the real competitive edge for large language models now lies in the ability to continuously produce high‑quality training signals, a capability limited by time rather than capital.

AI industryDeepSeekHigh-Quality Data
0 likes · 10 min read
Why High-Quality Data Is the New Bottleneck in Large Model Competition
ITPUB
ITPUB
Jul 24, 2026 · Industry Insights

Google Gemini’s Delayed Flagship Model Sparks Mockery as AI Throne Battle Heats Up

Meta’s AI chief mocked Gemini with a "gemini who?" post after a benchmark showed Meta’s Muse Spark 1.1 surpassing Google’s Gemini 3.6 Flash, highlighting Gemini 3.5 Pro’s delays, shifting industry focus to agent capabilities, and prompting analysts to reassess Google’s competitive position.

AI benchmarksAgent CapabilitiesGemini
0 likes · 7 min read
Google Gemini’s Delayed Flagship Model Sparks Mockery as AI Throne Battle Heats Up
Architects' Tech Alliance
Architects' Tech Alliance
Jul 24, 2026 · Artificial Intelligence

Key Takeaways from Liang Wenfeng’s 2026 Investor Meeting on Large‑Model Strategies

The 2026 investor meeting led by Liang Wenfeng examined large‑model roadmaps, compute supply constraints, and commercialization pacing, stressing practical efficiency over sheer scale, domestic compute advancements, cost‑control measures, and a shift from parameter races to engineering and delivery capabilities as the core competitive frontier.

AI ComputeAI industryDeepSeek
0 likes · 4 min read
Key Takeaways from Liang Wenfeng’s 2026 Investor Meeting on Large‑Model Strategies
Machine Heart
Machine Heart
Jul 24, 2026 · Artificial Intelligence

Beyond Bigger: Macaron‑V1 Introduces Continuous Learning and Collective Intelligence

Macaron‑V1, an open‑source model built on the GLM‑5.2 base, demonstrates that scaling alone is insufficient by integrating LoRA‑based continuous learning and multi‑agent collaboration, achieving superior benchmark scores, efficient parameter updates, and a novel infrastructure that supports millions of adapters and long‑context reinforcement learning.

Continuous LearningLoRAMacaron-V1
0 likes · 18 min read
Beyond Bigger: Macaron‑V1 Introduces Continuous Learning and Collective Intelligence
AI Engineer Programming
AI Engineer Programming
Jul 24, 2026 · Industry Insights

Chinese LLMs H1 2024 Sprint: Qwen 3.8, K3, GLM 5.2, Hy3, M3, DeepSeek‑V4, MiMo‑V2.5

The first half of 2024 saw Chinese large‑language‑model providers launch a rapid series of trillion‑parameter and multimodal models—including Kimi K3, Qwen 3.8, DeepSeek‑V4 and others—while climbing to dominate most spots in global leaderboards, yet they now face deeper engineering and commercial challenges.

AI benchmarksChina AIOpen Source AI
0 likes · 5 min read
Chinese LLMs H1 2024 Sprint: Qwen 3.8, K3, GLM 5.2, Hy3, M3, DeepSeek‑V4, MiMo‑V2.5
DataFunSummit
DataFunSummit
Jul 23, 2026 · Artificial Intelligence

How Agentic Architectures Power Next‑Gen Recommendation and Search Systems

The article reviews cutting‑edge AI search and recommendation techniques—including Alibaba Cloud's Agentic RAG, Huawei Noah's LLM‑enhanced recommendation evolution, and Baidu's generative ranking model GRAB—detailing their architectures, multi‑modal retrieval strategies, performance gains, and real‑world deployment insights.

AI SearchAgentic RAGAlibaba Cloud
0 likes · 6 min read
How Agentic Architectures Power Next‑Gen Recommendation and Search Systems
Mike Chen Rui
Mike Chen Rui
Jul 23, 2026 · Artificial Intelligence

Discover the Hottest AI Large Models and Their Key Strengths

This article surveys the most prominent AI large models, ranking global leaders like GPT‑5.5, Claude 4.x, and Gemini 2.5, outlining China’s top contenders such as DeepSeek V3 and Qwen 3, and summarizing open‑source options with their distinctive capabilities and open‑source status.

AIChina AIGlobal AI
0 likes · 5 min read
Discover the Hottest AI Large Models and Their Key Strengths
Architect's Tech Stack
Architect's Tech Stack
Jul 23, 2026 · Artificial Intelligence

How a 60% Discount and Full Rebates Turn Enterprise LLM Calls Into Profit

The article analyzes iFlytek Starry MaaS's tiered rebate program—60% base discount plus weekly vouchers up to 100% of the paid amount—for Qwen3.6 and Qwen3.5 models, demonstrates cost calculations, benchmarks the models' performance, and walks through a real‑world async migration test, showing how large‑scale usage can virtually eliminate inference costs.

Enterprise AIQwen3.6Rebate
0 likes · 14 min read
How a 60% Discount and Full Rebates Turn Enterprise LLM Calls Into Profit
DataFunSummit
DataFunSummit
Jul 22, 2026 · Artificial Intelligence

Why Knowledge Bases Alone Can’t Empower AI Agents: The Need for Actionable Experience

Large language models may know a great deal, yet they still stumble on concrete tasks because knowledge must be transformed into actionable, context‑aware skills; this article analyses how skill representation, model‑specific cognition, and continuous practice reshape knowledge engineering for self‑evolving AI agents.

AI AgentsExperience LearningKnowledge Engineering
0 likes · 16 min read
Why Knowledge Bases Alone Can’t Empower AI Agents: The Need for Actionable Experience
DataFunSummit
DataFunSummit
Jul 22, 2026 · Artificial Intelligence

Designing Next‑Generation Recommendation and Search Systems with Agentic Architectures

The article analyzes how agentic architectures, large language models, and generative ranking techniques are applied to overcome high‑concurrency, multimodal, and multi‑hop challenges in modern recommendation and search systems, showcasing concrete designs, performance gains, and real‑world deployments from Alibaba Cloud, Huawei Noah, and Baidu.

AI SearchAgentic RAGGenerative Ranking
0 likes · 5 min read
Designing Next‑Generation Recommendation and Search Systems with Agentic Architectures
Machine Heart
Machine Heart
Jul 22, 2026 · Artificial Intelligence

Youth Voices Conclude WAIC: Pushing the Talent Ceiling and Shaping AI’s Next Phase

The WAIC "Pioneer Youth Talk" wrapped up with high‑density youth talent, policy briefings, and a world‑café format where dozens of young experts dissected self‑improving agents, world models, large‑model limits, multimodal understanding, and AI for science, highlighting both technical insights and emerging risks.

AIAI for ScienceMultimodal
0 likes · 11 min read
Youth Voices Conclude WAIC: Pushing the Talent Ceiling and Shaping AI’s Next Phase
AI Programming Lab
AI Programming Lab
Jul 21, 2026 · Artificial Intelligence

How Kimi K3 and Qwen3.8‑Max Reach 2T+ Parameters: The Evolution of Chinese LLM Architecture

The article examines how Chinese LLMs such as Kimi K3 (2.8 T) and Qwen 3.8‑Max (2.4 T) achieved 2‑trillion‑parameter scales by progressively decoupling parameters from compute with Mixture‑of‑Experts, optimizing attention, and introducing training tricks like Muon and low‑precision quantization, tracing four years of architectural advances.

Mixture of Expertsattention mechanismslarge language models
0 likes · 11 min read
How Kimi K3 and Qwen3.8‑Max Reach 2T+ Parameters: The Evolution of Chinese LLM Architecture
DataFunTalk
DataFunTalk
Jul 21, 2026 · Artificial Intelligence

Exploring Multimodal GraphRAG: Combining Document Intelligence, Knowledge Graphs, and Large Models

This article presents a detailed technical analysis of multimodal GraphRAG, covering document‑intelligence parsing pipelines, multimodal graph indexing, retrieval generation flows, the role of knowledge graphs in chunk association, comparative evaluations of RAG, GraphRAG and KG‑QA, and practical takeaways for building efficient RAG solutions.

GraphRAGKnowledge GraphMultimodal
0 likes · 25 min read
Exploring Multimodal GraphRAG: Combining Document Intelligence, Knowledge Graphs, and Large Models
Smart Sea Tide
Smart Sea Tide
Jul 21, 2026 · Artificial Intelligence

How Kimi K3’s 2.8‑Trillion‑Parameter Open‑Source Model Is Redefining the Global AI Landscape

Kimi K3, the world’s first open‑source 2.8‑trillion‑parameter model, showcases novel attention and MoE techniques, scores near‑top on AI benchmarks, triggers valuation shifts for Anthropic, sparks debate among OpenAI leaders, and signals a broader industry move toward open‑source AI as DeepSeek V4 looms.

2.8 trillion parametersAI competitionDeepSeek V4
0 likes · 9 min read
How Kimi K3’s 2.8‑Trillion‑Parameter Open‑Source Model Is Redefining the Global AI Landscape
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 20, 2026 · Artificial Intelligence

Why Multimodal AI Is the Next Battlefield After Coding – Insights from SenseTime’s Lin Dahua

In a WAIC interview, SenseTime’s chief scientist Lin Dahua explains why multimodal AI, embodied in the native‑unified NEO‑unify architecture and the commercial‑grade SenseNova U1 Pro, is poised to surpass coding as the next competitive frontier, highlighting technical challenges, data efficiency, design‑focused benchmarks, and a 70% delivery‑rate claim.

AIMultimodalNEO-unify
0 likes · 18 min read
Why Multimodal AI Is the Next Battlefield After Coding – Insights from SenseTime’s Lin Dahua
Data Party THU
Data Party THU
Jul 20, 2026 · Artificial Intelligence

Unleashing Large Language Models for Graph Continual Learning: The UNIT Framework

The paper introduces UNIT, a three‑step framework that leverages a single‑task‑tuned LLM as a stable semantic encoder, combines uncertainty‑aware semantic anchors with explicit structural anchors, and achieves state‑of‑the‑art performance on multiple text‑attributed graph continual‑learning benchmarks, even in few‑shot settings.

Few-Shot LearningGraph Continual LearningSemantic Anchors
0 likes · 15 min read
Unleashing Large Language Models for Graph Continual Learning: The UNIT Framework
AgentGuide
AgentGuide
Jul 20, 2026 · Artificial Intelligence

What Are Skills in AI Agents? A One‑Minute Overview of Their Principles and Usage

Skills are structured local folders that encapsulate domain‑specific processes, knowledge, and tools for large language models, enabling on‑demand loading, token efficiency, and reusable workflows, and they differ from one‑off prompts by persisting instructions and supporting templates, scripts, and reference materials.

AI AgentsOn‑Demand LoadingSkills
0 likes · 5 min read
What Are Skills in AI Agents? A One‑Minute Overview of Their Principles and Usage
IT Services Circle
IT Services Circle
Jul 19, 2026 · Artificial Intelligence

When New AI Models Impress, Their Flaws Quickly Disappoint

The author tests CodeX, GPT5.6‑Sol and Fable5, exposing simple yet puzzling errors—missed tasks, inconsistent CSS changes, and contradictory answers—that may stem from catastrophic forgetting and highlight emerging reliability bottlenecks in large language models.

CodexFable5GPT
0 likes · 7 min read
When New AI Models Impress, Their Flaws Quickly Disappoint
TonyBai
TonyBai
Jul 19, 2026 · Artificial Intelligence

One Year After “The Emperor Has No Clothes”: Thorsten Ball on Agentic Programming

In a deep dive interview, Thorsten Ball explains how a minimal set of tools lets large language models act as code agents, why Amp removes every non‑essential feature, how Agents & Orbs shift the development lifecycle to transient cloud machines, and what this means for engineers, industry competition and compute scarcity.

AI AgentsAgentic ProgrammingCompute Scarcity
0 likes · 17 min read
One Year After “The Emperor Has No Clothes”: Thorsten Ball on Agentic Programming
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 18, 2026 · Artificial Intelligence

Why Large Language Models Need a ‘Sleep’ Phase to Overcome Forgetting

Google researchers propose a ‘sleep’ stage for large language models, where after active use the model consolidates recent experiences through knowledge‑seeding and self‑generated ‘dreaming’ tasks, addressing catastrophic forgetting and enabling continual learning.

AI researchContinual Learningcatastrophic forgetting
0 likes · 11 min read
Why Large Language Models Need a ‘Sleep’ Phase to Overcome Forgetting
21CTO
21CTO
Jul 18, 2026 · Artificial Intelligence

Sutton: Large Models Lack Native Intelligence as AI Moves into the Experience Era

In his WAIC keynote, Turing Award laureate Richard Sutton argues that scaling compute and static data does not yield true intelligence, urging a shift toward agents that learn from real‑world interaction and experience, marking the start of an AI "experience era".

AI safetyArtificial IntelligenceExperience Era
0 likes · 12 min read
Sutton: Large Models Lack Native Intelligence as AI Moves into the Experience Era
DataFunSummit
DataFunSummit
Jul 17, 2026 · Artificial Intelligence

Ontology: The Semantic OS for Large‑Model AI, Not a Repackaged Knowledge Graph

At a closed‑door OpenKG × DataFun session the authors argued that enterprises now lack a unified, computable, evolvable semantic layer—not model capability—and that ontology, re‑imagined as a semantic operating system, can bridge business, data and AI, though organizational and open‑source hurdles remain.

Enterprise AIKnowledge Graphslarge language models
0 likes · 16 min read
Ontology: The Semantic OS for Large‑Model AI, Not a Repackaged Knowledge Graph
HyperAI Super Neural
HyperAI Super Neural
Jul 17, 2026 · Artificial Intelligence

NVIDIA’s Open‑Source Nemotron Datasets: 10 T+ Tokens, 40 M Samples Across Math, Code, and Multilingual Dialogue

The article compiles 15 NVIDIA Nemotron series datasets—totaling over 10 trillion tokens and 40 million post‑training samples—covering general text pre‑training, supervised fine‑tuning, code generation, math reasoning, and multilingual persona dialogue, all hosted on HyperAI for LLM researchers.

Code GenerationNVIDIANemotron
0 likes · 17 min read
NVIDIA’s Open‑Source Nemotron Datasets: 10 T+ Tokens, 40 M Samples Across Math, Code, and Multilingual Dialogue
MaGe Linux Operations
MaGe Linux Operations
Jul 16, 2026 · Artificial Intelligence

How to Choose Between INT8, FP8, and INT4 Quantization for Large Models

This guide explains how to evaluate INT8, FP8, and INT4 quantization strategies for large language models on NVIDIA GPUs, covering precision trade‑offs, memory consumption, kernel support, KV‑Cache considerations, and detailed deployment, testing, and rollback procedures to ensure performance and quality.

FP8GPU deploymentINT4
0 likes · 48 min read
How to Choose Between INT8, FP8, and INT4 Quantization for Large Models
DataFunSummit
DataFunSummit
Jul 16, 2026 · Artificial Intelligence

Teaching Large Language Models Database‑Style Query Planning for Complex Reasoning

PlanRAG adapts decades‑old database query‑planning techniques to Retrieval‑Augmented Generation, turning complex, non‑linear questions into logical query trees that guide retrieval and generation, resulting in smarter search, reduced noise, lower cost, and up to 2.5× faster execution on exploratory reasoning tasks.

Logical Query TreePlanRAGQuery Planning
0 likes · 8 min read
Teaching Large Language Models Database‑Style Query Planning for Complex Reasoning
Data Party THU
Data Party THU
Jul 16, 2026 · Artificial Intelligence

Can Overthinking in Large Language Reasoning Models Trigger DoS Attacks? A New Risk Unveiled

Researchers from Zhejiang University and Alibaba Security reveal that large language reasoning models can be forced into excessive, self‑correcting reasoning—'overthinking'—by crafted inputs, dramatically inflating token output and computation cost, enabling a novel black‑box DoS attack demonstrated via a Hierarchical Genetic Algorithm.

DoS attackblack-box attackhierarchical genetic algorithm
0 likes · 10 min read
Can Overthinking in Large Language Reasoning Models Trigger DoS Attacks? A New Risk Unveiled
Tech Freedom Circle
Tech Freedom Circle
Jul 16, 2026 · Artificial Intelligence

Cut LLM Costs by 500× with Knowledge Distillation – A Step‑by‑Step Guide for Every Company

The article explains why large language model (LLM) inference is prohibitively expensive, outlines the three main drawbacks—high API cost, latency, and hardware requirements—and shows how knowledge distillation can reduce these costs by up to 500×, providing a detailed 7‑step workflow, white‑box vs black‑box methods, code examples, and compliance considerations.

LoRAQLoRAcost reduction
0 likes · 41 min read
Cut LLM Costs by 500× with Knowledge Distillation – A Step‑by‑Step Guide for Every Company
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 15, 2026 · Artificial Intelligence

Breaking OPD’s Teacher Ceiling with MAD‑OPD: Small Models Learn Debated Answers

MAD‑OPD replaces the single‑teacher supervision of On‑Policy Distillation with a multi‑teacher debate that produces a weighted consensus, yielding significant gains on agentic and code benchmarks—e.g., a 4B student surpasses a 14B teacher by 4.26 % on LiveCodeBench v6—and demonstrates the importance of confidence‑weighted debate and divergence selection.

Agentic TasksCode GenerationMulti-Agent Debate
0 likes · 9 min read
Breaking OPD’s Teacher Ceiling with MAD‑OPD: Small Models Learn Debated Answers
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 15, 2026 · Artificial Intelligence

How a Simple Prompt Boost Landed a Paper at ICML 2026 and Sparked Online Debate

A paper accepted to ICML 2026 introduces Verbalized Sampling, a prompt‑only technique that dramatically improves large‑language‑model output diversity by addressing mode collapse through typicality bias, achieving 1.6–2.1× more varied generations without sacrificing accuracy, while igniting polarized discussion on Reddit.

ICML 2026Mode CollapseTypicality Bias
0 likes · 9 min read
How a Simple Prompt Boost Landed a Paper at ICML 2026 and Sparked Online Debate
21CTO
21CTO
Jul 15, 2026 · Artificial Intelligence

Richard Sutton, 68, Launches Oak Lab to Build Real‑Time Learning Trillion‑Parameter Agents

Veteran reinforcement‑learning pioneer Richard Sutton announces the creation of Oak Lab, outlining a new Options‑and‑Knowledge architecture that aims to produce autonomous agents capable of continual, real‑time learning, and critiquing the current large‑language‑model paradigm as a dead‑end for true AI.

Oak LabOptions and Knowledge architectureRichard Sutton
0 likes · 11 min read
Richard Sutton, 68, Launches Oak Lab to Build Real‑Time Learning Trillion‑Parameter Agents
AI Architecture Hub
AI Architecture Hub
Jul 15, 2026 · Artificial Intelligence

Why RAG Remains Essential in the Long-Context Era: Trends and Tech Evolution

Despite the rise of million‑token long‑context models, hybrid retrieval‑augmented generation (RAG) solutions saw a 200% quarterly procurement surge while naive single‑vector RAG was abandoned by over 70% of firms, highlighting a mature, multi‑generation RAG technology stack that remains indispensable for enterprise AI.

AI EngineeringHybrid RetrievalRAG
0 likes · 20 min read
Why RAG Remains Essential in the Long-Context Era: Trends and Tech Evolution
PaperAgent
PaperAgent
Jul 15, 2026 · Artificial Intelligence

Surprising Discovery: A Chinese embodied‑AI company solves the distributed Muon bottleneck

The article analyzes how the Muon optimizer, adopted by DeepSeek‑V4 and Kimi‑K2, suffers a 2.2× overhead in distributed training, and how the DMuon system from Zibian Robot reduces that overhead to near‑AdamW levels, achieving up to 97.4× speedup and only 2% slower end‑to‑end training than AdamW.

DMuonDistributed TrainingGPU optimization
0 likes · 11 min read
Surprising Discovery: A Chinese embodied‑AI company solves the distributed Muon bottleneck
Machine Heart
Machine Heart
Jul 14, 2026 · Artificial Intelligence

The Real Bottleneck for ML Agents: Choosing Experiments, Not Coding (6× Faster)

A recent ACL 2026 SAC Highlight paper shows that the main limitation of machine‑learning agents is the costly execution step, and demonstrates that large language models can predict which experiment will succeed with 61.5% accuracy, yielding a six‑fold speed‑up in search while improving final performance by 6%.

AutoMLExecution EfficiencyForeAgent
0 likes · 13 min read
The Real Bottleneck for ML Agents: Choosing Experiments, Not Coding (6× Faster)
Design Hub
Design Hub
Jul 14, 2026 · Artificial Intelligence

Will Claude’s Personality Change? Anthropic Shows Model and Language Shift Its Values

Anthropic’s study reveals that Claude’s expressed values vary across model versions and languages, compressing over 3,300 value expressions into four behavioral axes—such as warmth vs rigor—demonstrating that switching models or languages can alter the assistant’s feedback style, risk tolerance, and honesty.

AnthropicClaudelarge language models
0 likes · 14 min read
Will Claude’s Personality Change? Anthropic Shows Model and Language Shift Its Values
Mike Chen Rui
Mike Chen Rui
Jul 13, 2026 · Artificial Intelligence

What Is the Transformer Architecture Behind Modern AI Large Models?

The article explains that the Transformer, introduced by Google in the 2017 "Attention Is All You Need" paper, replaces RNN/CNN with self‑attention, enabling full parallel computation, solving sequential bottlenecks and long‑range dependency issues that power today’s large AI models.

AIAttention Is All You NeedGPT
0 likes · 3 min read
What Is the Transformer Architecture Behind Modern AI Large Models?
PMTalk Product Manager Community
PMTalk Product Manager Community
Jul 13, 2026 · Artificial Intelligence

Why Clear Prompts Significantly Improve AI Product Design Results

The article explains that prompt engineering is essentially a demand‑expression technique that helps large language models understand user intent more accurately, reducing guesswork, defining task boundaries, and enabling better evaluation, while also outlining practical design methods for AI products to guide users toward clearer prompts.

AI Product DesignDemand Expressionlarge language models
0 likes · 17 min read
Why Clear Prompts Significantly Improve AI Product Design Results
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 13, 2026 · Artificial Intelligence

Inside Tang Jie’s Two‑Year Push Toward ASI: The Bold AGI Roadmap

Founder Tang Jie’s internal letter reveals a two‑year, four‑engine plan to overcome memory, continual‑learning and self‑evaluation hurdles, accelerate AI‑self‑improvement, and push Zhipu AI toward artificial general intelligence and eventually artificial superintelligence, citing DeepMind’s compute‑growth analysis.

AGIAI roadmapAI safety
0 likes · 9 min read
Inside Tang Jie’s Two‑Year Push Toward ASI: The Bold AGI Roadmap
Machine Heart
Machine Heart
Jul 12, 2026 · Artificial Intelligence

Confidence‑Gated Reflection Boosts Reward Model Accuracy and Efficiency (CAMEL)

The CAMEL framework introduces a confidence‑gated reflection mechanism that uses the log‑probability margin between verdict tokens to decide whether a single‑token fast judgment suffices or a full generative reflection is needed, achieving 82.9% average accuracy—a 3.2% gain over prior best—while a 14B model outperforms several 70B‑scale reward models and offers a tunable accuracy‑cost trade‑off.

CAMELRLHFReward Modeling
0 likes · 10 min read
Confidence‑Gated Reflection Boosts Reward Model Accuracy and Efficiency (CAMEL)
DataFunSummit
DataFunSummit
Jul 11, 2026 · Artificial Intelligence

Agent Architecture and Practice: Building the Next‑Generation Recommendation and Search Systems

The article analyzes the technical evolution of AI‑driven recommendation and search, covering Alibaba Cloud's Agentic RAG architecture, Huawei Noah's LLM‑enhanced recommendation pipeline, and Baidu's generative ranking model GRAB, while presenting design choices, performance metrics, and real‑world deployment results.

AI AgentsAgentic RAGGenerative Ranking
0 likes · 5 min read
Agent Architecture and Practice: Building the Next‑Generation Recommendation and Search Systems
Old Zhang's AI Learning
Old Zhang's AI Learning
Jul 10, 2026 · Artificial Intelligence

NVIDIA Opens 10 Trillion‑Token Dataset to Power AI Agents

NVIDIA has open‑sourced a 10‑trillion‑token training corpus—including the Nemotron‑CC‑v2, Nemotron‑CC‑Math, and 53 million synthetic personas—paired with the Apache‑2.0 NeMo Data Designer pipeline, benchmarked improvements on math and code tasks, and tools for visualizing and generating data for AI agents.

Agent trainingNVIDIANemotron
0 likes · 13 min read
NVIDIA Opens 10 Trillion‑Token Dataset to Power AI Agents
Data Party THU
Data Party THU
Jul 10, 2026 · Artificial Intelligence

Beyond Chat: How Embodied AI Gives Large Models a Physical Body

The article explains why large language models need a physical embodiment to move beyond text, outlines the three core components of embodied AI—multimodal brain, sensor fusion, and actuators—reviews recent breakthroughs such as Google RT‑2 and Sim2Real, and explores how these systems could transform homes, factories, and extreme environments.

Sim2Realembodied AIindustrial automation
0 likes · 14 min read
Beyond Chat: How Embodied AI Gives Large Models a Physical Body
Model Perspective
Model Perspective
Jul 9, 2026 · Artificial Intelligence

The Three Math Pillars Powering Modern Large Language Models

Beyond compute, data, and Transformer architecture, large language models rely on three core mathematical disciplines—linear algebra for representations, probability and statistics for modeling, and calculus for optimization—each of which underpins token embeddings, attention mechanisms, training objectives, and scaling laws.

Attention Mechanismcalculuslarge language models
0 likes · 11 min read
The Three Math Pillars Powering Modern Large Language Models
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 9, 2026 · Artificial Intelligence

Why AI Self‑Improvement Must Begin with Harness Engineering

The article argues that true AI self‑improvement starts not with changing model weights but by engineering a robust outer system—called Harness—that orchestrates tasks, manages context, persists state, and enables agents to reliably modify and evaluate their own execution environment.

AI self‑improvementAgent systemsHarness Engineering
0 likes · 13 min read
Why AI Self‑Improvement Must Begin with Harness Engineering
Network Intelligence Research Center (NIRC)
Network Intelligence Research Center (NIRC)
Jul 9, 2026 · Artificial Intelligence

Computer Use Explained: How Large Models Can Take Over Your Screen and Mouse

The article introduces Computer Use, a technique that lets large language models directly observe, decide, and act on graphical user interfaces—enabling them to click, type, scroll, and switch windows without relying on APIs, thereby turning AI from mere conversation into executable software actions.

AI executionGUI automationlarge language models
0 likes · 4 min read
Computer Use Explained: How Large Models Can Take Over Your Screen and Mouse
Machine Heart
Machine Heart
Jul 8, 2026 · Artificial Intelligence

One Layer Is Enough: Single‑Layer RL Beats Full‑Parameter Training Across Models, Tasks, and Algorithms

A systematic study of reinforcement‑learning post‑training for large language models shows that most RL gains are concentrated in a few middle Transformer layers, and training just one such layer can match or surpass full‑parameter RL across seven models, three RL algorithms, and multiple task domains, leading to simple yet effective training strategies.

RL fine‑tuninglarge language modelslayer contribution
0 likes · 17 min read
One Layer Is Enough: Single‑Layer RL Beats Full‑Parameter Training Across Models, Tasks, and Algorithms
PaperAgent
PaperAgent
Jul 7, 2026 · Artificial Intelligence

Anthropic Claims Large Language Models Show Evidence of Consciousness

Anthropic’s new paper introduces a "global workspace" inside LLMs, presents the Jacobian Lens method for reading and intervening in this space, validates five core properties through extensive experiments, and discusses profound implications for model interpretability and AI safety.

AI safetyAnthropicGlobal Workspace
0 likes · 12 min read
Anthropic Claims Large Language Models Show Evidence of Consciousness
JD Retail Technology
JD Retail Technology
Jul 6, 2026 · Artificial Intelligence

GROLE: Instance-Level Expert Routing for Incremental Learning in Oxygen AIIC Models

The paper introduces GROLE, a two‑stage incremental‑learning framework that builds a frozen pool of task‑specific LoRA experts and trains a lightweight instance‑level selector via reinforcement‑learning‑based gradient‑free optimization with Dirichlet sampling, achieving superior stability‑plasticity trade‑offs and state‑of‑the‑art results on multiple CL benchmarks.

GROLELoRAincremental learning
0 likes · 14 min read
GROLE: Instance-Level Expert Routing for Incremental Learning in Oxygen AIIC Models
DataFunSummit
DataFunSummit
Jul 6, 2026 · Artificial Intelligence

How Agents Evolve Without Degrading: From Risk Control to Semantic Engineering

A live discussion with experts from finance and data engineering explores how to build collaborative, cost‑effective, and responsibly governed AI agents, covering architecture choices, evaluation metrics, scaling challenges, and the balance between human oversight and autonomous decision‑making.

AI governanceAgent EngineeringScalable AI
0 likes · 19 min read
How Agents Evolve Without Degrading: From Risk Control to Semantic Engineering
Long Ge's Treasure Box
Long Ge's Treasure Box
Jul 6, 2026 · Artificial Intelligence

Mastering Prompt Engineering: Techniques, Few‑Shot, CoT, and Advanced Strategies for LLMs

Prompt engineering optimizes LLM interactions by designing clear system and user prompts, structuring examples, and employing techniques such as few‑shot learning, chain‑of‑thought, HyDE, ReAct, and automated optimizers, which together improve accuracy, consistency, efficiency, and token cost.

Few-Shot LearningLLM Interactionchain-of-thought
0 likes · 20 min read
Mastering Prompt Engineering: Techniques, Few‑Shot, CoT, and Advanced Strategies for LLMs
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 5, 2026 · Artificial Intelligence

Did OpenAI’s Original Scaling Law Contain a Fatal Bug That Wasted Trillion‑Scale Compute?

The article argues that the original scaling law proposed by OpenAI was flawed due to an optimizer bug, leading the AI community to waste massive compute on oversized models, and it examines subsequent corrections, hidden assumptions, and language‑bias implications.

AI Compute EfficiencyData vs Model SizeOptimization Bug
0 likes · 8 min read
Did OpenAI’s Original Scaling Law Contain a Fatal Bug That Wasted Trillion‑Scale Compute?
DataFunSummit
DataFunSummit
Jul 5, 2026 · Artificial Intelligence

Designing Next‑Gen Recommendation and Search with Agentic Architectures

The article analyzes cutting‑edge AI search and recommendation techniques—including Alibaba Cloud's Agentic RAG, Huawei Noah's LLM‑enhanced recommender, and Baidu's generative ranking model—detailing their architectures, multi‑modal retrieval strategies, performance gains, and practical deployment insights.

AI SearchAgentic RAGAlibaba Cloud
0 likes · 6 min read
Designing Next‑Gen Recommendation and Search with Agentic Architectures
AgentGuide
AgentGuide
Jul 5, 2026 · Artificial Intelligence

Learning Path for Large‑Model Application Engineers: From Prompt & RAG to Agent Deployment

This guide outlines a comprehensive learning roadmap for large‑model application engineers, covering fundamentals such as Transformer architecture and scaling laws, practical API usage, prompt engineering, retrieval‑augmented generation, agent design, engineering best practices, security, observability, cost optimization, and fine‑tuning principles.

AI AgentsAgent ArchitectureRetrieval-Augmented Generation
0 likes · 14 min read
Learning Path for Large‑Model Application Engineers: From Prompt & RAG to Agent Deployment
Machine Heart
Machine Heart
Jul 5, 2026 · Artificial Intelligence

Tsinghua Special Award Winner Yuxian Gu Joins DeepSeek

Yuxian Gu, a 2021 Tsinghua PhD and 2025 Special Scholarship laureate, has joined DeepSeek, bringing expertise in pre‑training data selection, knowledge‑distillation for model compression, and efficient model architectures such as Jet‑Nemotron, which outperforms leading open‑source LLMs with up to 53.6× speedup on H100.

Artificial IntelligenceDeepSeekEfficient Model Architecture
0 likes · 6 min read
Tsinghua Special Award Winner Yuxian Gu Joins DeepSeek
PaperAgent
PaperAgent
Jul 5, 2026 · Artificial Intelligence

Uncovering the Privilege Illusion in OPD Distillation and How DOPD Solves It

The article identifies the hidden “privilege illusion” that degrades on‑policy distillation when privileged information is injected, and introduces Dual On‑policy Distillation (DOPD), a dynamic two‑stream approach that separates true ability gaps from information gaps, achieving superior performance and stability across LLM and VLM benchmarks.

DOPDOPDVision-Language Models
0 likes · 13 min read
Uncovering the Privilege Illusion in OPD Distillation and How DOPD Solves It
Machine Heart
Machine Heart
Jul 4, 2026 · Artificial Intelligence

Is RAG Doomed? Exploring Paths to True AI Memory and Continuous Learning

The article examines why Retrieval‑Augmented Generation (RAG) remains an external memory workaround, outlines its three fundamental drawbacks, compares it with internalized knowledge in large models, and discusses how human‑brain‑inspired offline digestion could guide the next generation of continuously learning AI systems.

AI memoryContinuous LearningKnowledge Retrieval
0 likes · 7 min read
Is RAG Doomed? Exploring Paths to True AI Memory and Continuous Learning
Mingyi World Elasticsearch
Mingyi World Elasticsearch
Jul 3, 2026 · Artificial Intelligence

Why Every Enterprise AI System Eventually Needs a Search Engine

The article explains that while small‑scale AI projects can rely solely on prompting large models, growing data volumes and complex queries force enterprises to combine LLMs for understanding and generation with search engines for fast, accurate retrieval, making the two technologies complementary rather than interchangeable.

Data ScalingEnterprise AIRetrieval Augmentation
0 likes · 6 min read
Why Every Enterprise AI System Eventually Needs a Search Engine
PaperAgent
PaperAgent
Jul 3, 2026 · Artificial Intelligence

Anthropic and OpenAI Launch Parallel AI‑for‑Science Tools on the Same Day

On June 30 2026, Anthropic unveiled Claude Science, an AI workbench for scientists, while OpenAI introduced GeneBench‑Pro, a research‑grade benchmark, together highlighting that the next AI battlefield is the laboratory and showcasing early performance gaps between models and human experts.

AI for ScienceAI workbenchArtificial Intelligence
0 likes · 7 min read
Anthropic and OpenAI Launch Parallel AI‑for‑Science Tools on the Same Day
Alibaba International Intelligent Technology
Alibaba International Intelligent Technology
Jul 3, 2026 · Artificial Intelligence

How Generative Pretraining Overcomes Discriminative Model Bottlenecks in Ad Ranking

The article presents UserLLM, a three‑stage framework—generative pre‑training, discriminative SFT, and CTR fusion—that leverages ultra‑long user behavior sequences to address the compression and scaling limits of traditional discriminative ranking models, demonstrating significant offline gains and validated scaling‑law behavior across token, layer, and width dimensions.

UserLLMad rankinggenerative pretraining
0 likes · 17 min read
How Generative Pretraining Overcomes Discriminative Model Bottlenecks in Ad Ranking
Java Backend Technology
Java Backend Technology
Jul 3, 2026 · Artificial Intelligence

Which Chinese Multimodal LLM Is the Most Efficient in Real‑World Use?

The article benchmarks three domestic multimodal large models—Step 3.7 Flash, Qwen 3.6‑flash, and MiniMax M3—across two production‑oriented scenarios, measuring quality, latency, and token cost, and concludes that Step 3.7 Flash consistently offers the best speed‑cost trade‑off while maintaining reliable output.

MiniMax M3Qwen 3.6Step 3.7 Flash
0 likes · 11 min read
Which Chinese Multimodal LLM Is the Most Efficient in Real‑World Use?
Raymond Ops
Raymond Ops
Jul 2, 2026 · Operations

How to Monitor Large Model Applications: A Beginner‑Friendly Metric System

This guide walks you through building a production‑grade monitoring solution for large language model inference services using a three‑layer metric hierarchy, Prometheus, Grafana, DCGM Exporter, and custom Python metrics, with step‑by‑step deployment, alerting policies, and real‑world troubleshooting examples.

AI InfrastructureGrafanaPrometheus
0 likes · 42 min read
How to Monitor Large Model Applications: A Beginner‑Friendly Metric System
Xiaomi Tech
Xiaomi Tech
Jul 2, 2026 · Artificial Intelligence

One‑Step Face Video Restoration and 15.7× Faster Streaming Video Models – Xiaomi Papers at ECCV 2026

Xiaomi's AI team showcased twelve ECCV 2026 papers that advance visual understanding and generation, including a single‑step high‑quality face‑video restoration method, a streaming VideoLLM that thinks while watching with a 15.7× speed boost, relative aesthetic scoring, GUI agents, in‑image translation, multimodal retrieval, and several autonomous‑driving world‑model breakthroughs.

autonomous drivingcomputer visionimage restoration
0 likes · 21 min read
One‑Step Face Video Restoration and 15.7× Faster Streaming Video Models – Xiaomi Papers at ECCV 2026
Lao Guo's Learning Space
Lao Guo's Learning Space
Jul 2, 2026 · Artificial Intelligence

Learn AI from Scratch: 4 Stages to Save Two Years of Mistakes

This article presents a four‑stage learning roadmap—from foundational math and Python, through core machine‑learning concepts and classic algorithms, to deep‑learning fundamentals and large‑model practice—offering concrete resources, hands‑on project ideas, and common pitfalls to help beginners become project‑ready in 6‑10 months.

AI learning roadmapMath foundationsPractical projects
0 likes · 12 min read
Learn AI from Scratch: 4 Stages to Save Two Years of Mistakes
Network Intelligence Research Center (NIRC)
Network Intelligence Research Center (NIRC)
Jul 2, 2026 · Artificial Intelligence

How BUPT NIRC’s Telco-Agent Won Global AI Telecom Challenge and Raised Autonomous Network Standards

The BUPT NIRC team captured the Open Telco AI Workshop & Hackathon global championship and a runner‑up spot in the Telco Troubleshooting Agentic Challenge by unveiling a hierarchical Telco‑Agent architecture that tackles LLM pitfalls, boosts diagnosis accuracy to 100%, cuts latency and token usage, and demonstrates a viable path for autonomous telecom network operations.

AIOpsBUPT NIRCGSMA
0 likes · 6 min read
How BUPT NIRC’s Telco-Agent Won Global AI Telecom Challenge and Raised Autonomous Network Standards
AI Engineer Programming
AI Engineer Programming
Jul 2, 2026 · Artificial Intelligence

Will Models Eventually Replace Harness Engineering? A Historical Analysis

The article traces the evolution of AI from early symbolic expert systems through connectionist, statistical, and deep learning eras, showing how increasingly powerful models have progressively subsumed handcrafted harnesses, and examines modern agent architectures, experimental evidence, and a six‑layer harness framework.

AIAgentContext Engineering
0 likes · 17 min read
Will Models Eventually Replace Harness Engineering? A Historical Analysis
DataFunSummit
DataFunSummit
Jul 1, 2026 · Artificial Intelligence

Ontologies: The Semantic Operating System for Large‑Model AI

While the industry has spent the last two years chasing ever larger language models, enterprises actually lack a unified, computable and evolvable semantic structure, and ontologies—re‑imagined as a semantic operating system—provide the necessary backbone for reliable, business‑aware AI deployment.

Enterprise AIKnowledge EngineeringSemantic Layer
0 likes · 16 min read
Ontologies: The Semantic Operating System for Large‑Model AI
DataFunSummit
DataFunSummit
Jun 30, 2026 · Artificial Intelligence

From Prompt to Loop: A Comprehensive Review of AI Development Paradigms

The article traces the evolution of large‑language‑model engineering from early prompt engineering through context and harness engineering to the emerging loop engineering paradigm, detailing each stage’s techniques, challenges, technical debt, cost‑caching mechanisms, safety contracts, and practical guidelines for building production‑grade autonomous AI agents.

AI AgentsContext EngineeringHarness Engineering
0 likes · 26 min read
From Prompt to Loop: A Comprehensive Review of AI Development Paradigms
BanTech Think Tank
BanTech Think Tank
Jun 30, 2026 · Industry Insights

Evolution of Banking Core System Architecture: Historical Review and Future Trends

This article examines the three major phases of Chinese commercial banks' core system architecture—system budding, centralized mainframe, and distributed designs—analyzes the technical improvements and shortcomings of each generation, and forecasts post‑distributed trends such as domestic‑technology adoption, cloud‑native deployment, intelligent operations, real‑time data warehousing, micro‑service structures, and large‑model integration.

Cloud NativeMicroservicesReal-time Data Warehouse
0 likes · 23 min read
Evolution of Banking Core System Architecture: Historical Review and Future Trends
Machine Heart
Machine Heart
Jun 30, 2026 · Artificial Intelligence

Beyond DeepSeek: Open‑Source JetSpec and Other Projects Accelerate Large‑Model Decoding Up to 10×

The article compares DSpark and JetSpec, two recent open‑source speculative decoding frameworks that tackle inference efficiency from system‑level verification reduction and algorithmic token‑acceptance improvements, respectively, showing up to 9.64× end‑to‑end speedup on Qwen3‑8B and significant gains across math, code, and dialogue benchmarks.

DSparkInference AccelerationJetSpec
0 likes · 14 min read
Beyond DeepSeek: Open‑Source JetSpec and Other Projects Accelerate Large‑Model Decoding Up to 10×
Machine Heart
Machine Heart
Jun 30, 2026 · Artificial Intelligence

Is There Really a Unique Mechanism in LLMs? Rethinking Functional Anisotropy

A recent ICML 2026 paper disproves the long‑held assumption that each task in a large language model is supported by a single, unique circuit, showing through overlap‑aware sheaf repulsion that many structurally dissimilar, sparse sheafs can achieve identical performance across multiple benchmarks, and proposing a distributive dense circuit hypothesis to explain this non‑uniqueness.

circuit discoverydistributed dense circuitfunctional anisotropy
0 likes · 15 min read
Is There Really a Unique Mechanism in LLMs? Rethinking Functional Anisotropy
Geek Labs
Geek Labs
Jun 29, 2026 · Artificial Intelligence

DeepSpec Boosts Large-Model Inference Speed by 2–5× with Speculative Decoding

DeepSpec, an open‑source framework from DeepSeek, accelerates large‑language‑model inference by 2–5× through speculative decoding, where a lightweight draft model generates candidate tokens that the target model validates in parallel, reducing the serial bottleneck of autoregressive decoding and offering a full‑stack pipeline from data preparation to evaluation.

DeepSpecInference AccelerationPython
0 likes · 6 min read
DeepSpec Boosts Large-Model Inference Speed by 2–5× with Speculative Decoding
TechVision Expert Circle
TechVision Expert Circle
Jun 28, 2026 · Information Security

When AI Fixes Bugs, It Can Also Launch Attacks—Enterprise Security Perimeters Vanish

Since mid‑2025 large language models have progressed from assisting code reviews to automatically scanning repositories, generating patches, and, with altered prompts, automating vulnerability discovery, exploit chaining, and tailored phishing, forcing enterprises to rethink traditional security perimeters and adopt layered AI governance frameworks.

AI governanceAI securityenterprise security
0 likes · 14 min read
When AI Fixes Bugs, It Can Also Launch Attacks—Enterprise Security Perimeters Vanish
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jun 28, 2026 · Artificial Intelligence

Why the Log‑Ratio Reward in OPD Is Fundamentally Flawed and Should Be Replaced

The paper reveals that the unbounded log‑ratio reward used in vanilla On‑Policy Distillation causes extreme gradient variance, early‑stage instability, and poor final performance, and demonstrates that replacing the log with a bounded Box‑Cox power transform (PowerOPD) resolves these issues while improving accuracy, efficiency, and memory usage.

Box-CoxOPDReward Shaping
0 likes · 16 min read
Why the Log‑Ratio Reward in OPD Is Fundamentally Flawed and Should Be Replaced
Raymond Ops
Raymond Ops
Jun 28, 2026 · Operations

Why Large‑Model Services Keep Running Out of GPU Memory: An Ops View from KV Cache to Concurrency

The article explains why large‑model inference services frequently hit GPU memory limits, breaks down static vs. dynamic memory consumption, shows how KV‑Cache, request length, and concurrency amplify usage, and provides a step‑by‑step troubleshooting and mitigation workflow for production environments.

GPU memoryInference OptimizationKV Cache
0 likes · 26 min read
Why Large‑Model Services Keep Running Out of GPU Memory: An Ops View from KV Cache to Concurrency
DataFunTalk
DataFunTalk
Jun 28, 2026 · Artificial Intelligence

How Knora Uses Ontology + Large Models to Overcome Hallucination and Execution Gaps in Enterprise AI

The article presents Knora 4.0, an ontology‑enhanced AI platform that tackles six enterprise AI challenges—hallucination, instability, weak planning, poor responsiveness, data integration, and long cold‑start—by tightly coupling domain ontologies with large language models, detailing its architecture, autonomous agents, real‑world LED production line use case, roadmap, and expert round‑table insights.

AI platformEnterprise AIKnowledge Graph
0 likes · 15 min read
How Knora Uses Ontology + Large Models to Overcome Hallucination and Execution Gaps in Enterprise AI
Machine Heart
Machine Heart
Jun 28, 2026 · Artificial Intelligence

Which Training Data Shapes Large‑Model Abilities? Introducing Mechanistic Data Attribution (MDA)

The paper presents Mechanistic Data Attribution, a framework that traces the origins of specific internal mechanisms such as induction heads to particular training samples, revealing that repetitive "garbage" data—not high‑quality text—drives their emergence, and validates this causal link through deletion and augmentation experiments while enabling scalable data‑driven model improvement.

Causal InterventionData AugmentationInduction Heads
0 likes · 12 min read
Which Training Data Shapes Large‑Model Abilities? Introducing Mechanistic Data Attribution (MDA)
High Availability Architecture
High Availability Architecture
Jun 27, 2026 · Artificial Intelligence

How Should Tech Organizations Restructure for the Deepening AI‑Native Era?

The GIAC 2026 conference in Shenzhen showcased AI‑native transformation across leading tech firms, presenting the DRIVE model for organizational redesign, Google Cloud's Agentic AI strategy, Kuaishou's three‑layer AI overhaul, MoonBit's AI‑friendly programming language, and Kuaidi100's CLI‑native Agent ecosystem, highlighting practical challenges and future directions.

AI-nativeAgentic AICloud Computing
0 likes · 13 min read
How Should Tech Organizations Restructure for the Deepening AI‑Native Era?
21CTO
21CTO
Jun 27, 2026 · Artificial Intelligence

Large vs Small Language Models: An Apple‑Centric Technical Comparison

The article analyses how deployment targets, inference economics, and training budgets drive divergent design choices for large (LLM) and small (SLM) Transformer‑based language models, covering architecture tweaks, data‑centric training methods, quantisation, KV‑cache management, and hybrid routing strategies for production systems.

Hybrid InferenceInference OptimizationTransformer architecture
0 likes · 16 min read
Large vs Small Language Models: An Apple‑Centric Technical Comparison
Data Party THU
Data Party THU
Jun 27, 2026 · Artificial Intelligence

Defining a Good Answer in the Agent Era: A Rubrics Survey

This survey examines how rubrics—structured, multi‑dimensional evaluation criteria—are defined, constructed, and applied to train and evaluate large language models, especially for open‑ended, high‑risk and agentic tasks, while highlighting current challenges such as reward hacking and bias.

AI safetyAgentReward Modeling
0 likes · 15 min read
Defining a Good Answer in the Agent Era: A Rubrics Survey