Tagged articles

large context

7 articles · Page 1 of 1
Machine Heart
Machine Heart
Jul 28, 2026 · Artificial Intelligence

LLaDA2.2: The First Large‑Scale Agentic Diffusion Model and Its Breakthroughs

LLaDA2.2 introduces Levenshtein‑based edit operations, 128K native context, and block routing to turn diffusion language models into reliable agents, achieving competitive scores on 17 benchmarks, up to 1.64× higher throughput than comparable autoregressive models, and demonstrating a new path for agentic AI.

Agentic AILLaDA2.2Levenshtein editing
0 likes · 16 min read
LLaDA2.2: The First Large‑Scale Agentic Diffusion Model and Its Breakthroughs
AntTech
AntTech
Jul 27, 2026 · Artificial Intelligence

LLaDA2.2 Released: Levenshtein Editing Enables Diffusion Language Models to Correct On-the-Fly

LLaDA2.2 introduces a Levenshtein‑based edit mechanism and the L‑EBPO reinforcement framework, allowing diffusion language models to delete and insert tokens during agent interactions, achieving near‑autoregressive accuracy (53.83 vs 55.74) and 1.64× higher BF16 throughput, plus an 8.6 % gain on SWE‑bench.

LLaDA2.2Levenshtein editingMoE
0 likes · 11 min read
LLaDA2.2 Released: Levenshtein Editing Enables Diffusion Language Models to Correct On-the-Fly
AI Explorer
AI Explorer
May 7, 2026 · Artificial Intelligence

Nvidia Endorses Open-Source “Light-Speed” Inference Engine for Coding Agents

The article examines how Nvidia’s open-source ‘light-speed’ inference engine tackles the token-bloat and compute bottlenecks of modern coding agents by redesigning attention and memory management, enabling order-of-magnitude speed gains without losing accuracy, and reshaping the AI-as-a-service ecosystem.

AI inferenceAttention optimizationNVIDIA
0 likes · 6 min read
Nvidia Endorses Open-Source “Light-Speed” Inference Engine for Coding Agents
Data Party THU
Data Party THU
Mar 21, 2026 · Artificial Intelligence

Why Bigger Context Windows Hurt LLMs and How RAG Still Wins

The article explains that expanding LLM context windows leads to attention dilution and retrieval collapse, degrading answer quality, and argues that Retrieval‑Augmented Generation remains essential because it preserves signal density through focused retrieval and selective prompting.

AI ArchitectureAttention DilutionLLM
0 likes · 8 min read
Why Bigger Context Windows Hurt LLMs and How RAG Still Wins
ShiZhen AI
ShiZhen AI
Mar 6, 2026 · Artificial Intelligence

GPT-5.4 Beats Human Baseline and Cuts Agent Token Use by Half

OpenAI's newly released GPT-5.4 integrates reasoning, coding, computer use, and agent tool calls, achieving a 75% success rate on OSWorld-Verified tasks—surpassing the human baseline—while its Tool Search feature reduces agent token consumption by 47% and supports up to 1 million tokens for long‑running workflows.

AI modelAgentComputer Use
0 likes · 15 min read
GPT-5.4 Beats Human Baseline and Cuts Agent Token Use by Half
Smart Sea Tide
Smart Sea Tide
Dec 16, 2025 · Artificial Intelligence

Google’s Gemini Deep Research vs OpenAI’s GPT‑5.2: Same‑Day Launch Sparks AI Rivalry

On the same day, Google unveiled Gemini Deep Research, a low‑hallucination, citation‑rich research agent built on Gemini 3 Pro, while OpenAI released GPT‑5.2 with multimodal, massive‑context capabilities and three pricing tiers, highlighting a strategic split between vertical depth and horizontal generalization backed by benchmark results.

AI model comparisonGPT-5.2Gemini Deep Research
0 likes · 10 min read
Google’s Gemini Deep Research vs OpenAI’s GPT‑5.2: Same‑Day Launch Sparks AI Rivalry