Tagged articles

LLM

2670 articles · Page 2 of 27
PaperAgent
PaperAgent
Aug 13, 2026 · Artificial Intelligence

First Community Benchmarks of DeepSeek V4 Pro, Qwen 3.8 Max, and Grok 4.6

The community quickly tested three newly released LLMs—DeepSeek V4 Pro, Qwen 3.8 Max, and Grok 4.6—across 3D scene generation, Flappy game creation, and airplane‑animation tasks, comparing quality, speed, and cost to reveal each model’s strengths and trade‑offs.

AIDeepSeekGrok
0 likes · 5 min read
First Community Benchmarks of DeepSeek V4 Pro, Qwen 3.8 Max, and Grok 4.6
AliExpress Tech
AliExpress Tech
Aug 13, 2026 · Artificial Intelligence

Building a Full‑Chain AI Agent for Intelligent Ticketing: From Config to Closed‑Loop

The article details a full‑chain intelligent ticketing platform that lets engineers configure tools and skills, run AI agents to diagnose and resolve tickets automatically, and continuously improve through a closed‑loop feedback system, highlighting architecture, runtime flow, configuration steps, and operational metrics.

AI agentLLMReact
0 likes · 10 min read
Building a Full‑Chain AI Agent for Intelligent Ticketing: From Config to Closed‑Loop
SuanNi
SuanNi
Aug 13, 2026 · Artificial Intelligence

DeepSeek V4 Pro vs. Grok 4.6: How New LLMs Challenge Top Closed‑Source Models

The newly released DeepSeek V4 Pro and Elon Musk’s Grok 4.6 deliver performance and cost metrics that rival or surpass leading closed‑source LLMs, with DeepSeek achieving up to 29‑fold cheaper token output and top scores on Agent, CyberGym, AutomationBench, Terminal‑Bench, and professional legal benchmarks, while Grok 4.6 matches GPT‑5.6 on the AA Intelligence Index and leads in workplace knowledge tests.

AIDeepSeekGrok
0 likes · 6 min read
DeepSeek V4 Pro vs. Grok 4.6: How New LLMs Challenge Top Closed‑Source Models
Machine Heart
Machine Heart
Aug 12, 2026 · Information Security

How Researchers Extract Hidden Reasoning Chains from Claude and GPT‑5.6

A new security paper demonstrates that design flaws in Claude, GPT‑5.6 and other leading LLM APIs allow attackers to steal encrypted reasoning blocks, replay them in weaker compatible models, and reconstruct most of the hidden thought process, exposing privacy and safety risks.

ClaudeGPT-5.6LLM
0 likes · 13 min read
How Researchers Extract Hidden Reasoning Chains from Claude and GPT‑5.6
Architecture Digest
Architecture Digest
Aug 11, 2026 · Backend Development

Run Your First Embabel Java Agent in 30 Minutes: A Hands‑On Guide

This article walks you through setting up the environment, creating a Spring Boot project, defining strong‑typed domain models, implementing @Action methods, declaring goals, and running an interactive shell so you can build and execute a fully functional Embabel Java Agent that automatically generates a research brief.

AIEmbabelJava
0 likes · 9 min read
Run Your First Embabel Java Agent in 30 Minutes: A Hands‑On Guide
SuanNi
SuanNi
Aug 11, 2026 · Artificial Intelligence

How Meta’s Open‑Source 30B Muse Glimmer Agent Runs on Your PC

Meta’s newly open‑sourced 30‑billion‑parameter Muse Glimmer agent model runs on a single consumer‑grade GPU, outperforms Gemma‑4 and Qwen‑3.6 on multiple Agent benchmarks, uses a perception encoder for multimodal input, and fits into a 20 GB memory envelope through quantization and a lightweight drafter.

LLMMultimodalMuse Glimmer
0 likes · 7 min read
How Meta’s Open‑Source 30B Muse Glimmer Agent Runs on Your PC
21CTO
21CTO
Aug 11, 2026 · Artificial Intelligence

Meta’s Muse Glimmer Open‑Source Release Revives the Open‑Weight Llama Competition

Meta has unveiled Muse Glimmer, a 30‑billion‑parameter open‑source LLM under Apache 2.0, positioned for agent workloads and benchmarked against Google’s Gemma 4 and Alibaba’s Qwen, while highlighting hardware requirements, performance limits, and the broader strategic implications for U.S. AI policy.

LLMMetaMuse Glimmer
0 likes · 10 min read
Meta’s Muse Glimmer Open‑Source Release Revives the Open‑Weight Llama Competition
Architect
Architect
Aug 10, 2026 · Artificial Intelligence

Anthropic Deep Dive: Context Engineering Lessons from Real‑World R&D

The article analyzes Anthropic’s “Effective context engineering for AI agents,” showing how larger context windows can degrade, categorizing information by stability, designing prompts in the Goldilocks zone, structuring tool contracts, and applying runtime information scheduling, compression, structured notes, and sub‑agents to keep AI agents reliable in complex development workflows.

AI agentsAnthropicContext Engineering
0 likes · 19 min read
Anthropic Deep Dive: Context Engineering Lessons from Real‑World R&D
DataFunSummit
DataFunSummit
Aug 9, 2026 · Artificial Intelligence

From Flawed RAG to Production‑Ready: A Deep Dive into Scaling Retrieval‑Augmented Generation

The article analyses why early RAG deployments suffer from low recall, hallucinations and cost overruns, breaks down eight concrete pain points—from PDF parsing pitfalls to the lost‑in‑the‑middle effect—then presents a systematic diagnosis framework, proven best‑practice roadmap, advanced GraphRAG and Agentic RAG approaches, and practical engineering trade‑offs for enterprise rollout.

Agentic RAGGraphRAGHybrid Search
0 likes · 19 min read
From Flawed RAG to Production‑Ready: A Deep Dive into Scaling Retrieval‑Augmented Generation
Machine Heart
Machine Heart
Aug 9, 2026 · Artificial Intelligence

Why Continual Learning Won’t Take Ten Years—Five Hot Paths and the Fight Against Catastrophic Forgetting

The article surveys five emerging approaches to LLM continual learning—external agent memory, context engineering, post‑training, pre‑training, and self‑modifying models—explaining how each tackles the core obstacle of catastrophic forgetting, citing benchmarks such as TRACE, ACE, and SDFT, and reflecting on Karpathy’s ten‑year timeline.

Context EngineeringContinual LearningLLM
0 likes · 17 min read
Why Continual Learning Won’t Take Ten Years—Five Hot Paths and the Fight Against Catastrophic Forgetting
PaperAgent
PaperAgent
Aug 9, 2026 · Artificial Intelligence

Tsinghua Unveils Two Breakthrough Papers on LLM Agent Skills

The article reviews Tsinghua University's two new papers—GSE, which introduces a global skill‑relation graph, clustering, and replay verification to make agent skills continuously improve, and SkillSentry, which uses ability contracts and adaptive honey‑world testing to ensure skill safety—detailing their methods, experimental results, and practical implications.

AI safetyAgentGSE
0 likes · 8 min read
Tsinghua Unveils Two Breakthrough Papers on LLM Agent Skills
webdream
webdream
Aug 8, 2026 · Artificial Intelligence

Engineering a Multi‑Agent System: Architecture, Stability, and Observability Lessons

This article shares practical engineering insights from building a multi‑agent LLM system, covering why multiple agents are needed, the 3‑agent + 1 skill architecture, LangGraph orchestration, tool integration via MCP, stability mechanisms, layered memory, traceability, streaming UI, and common pitfalls.

LLMLangGraphMCP
0 likes · 12 min read
Engineering a Multi‑Agent System: Architecture, Stability, and Observability Lessons
21CTO
21CTO
Aug 8, 2026 · R&D Management

Rust Team Issues AI Coding Policy: Allow LLMs but Ban “Pseudo‑Effort Signals”

The Rust project introduced a detailed LLM usage policy that permits AI for analysis while forbidding AI‑generated code creation, outlines five mandatory rules for AI‑derived contributions, sets safety red lines for security‑critical changes, and explains the community’s reaction and broader governance implications.

AIEngineering ManagementLLM
0 likes · 6 min read
Rust Team Issues AI Coding Policy: Allow LLMs but Ban “Pseudo‑Effort Signals”
Data Party THU
Data Party THU
Aug 8, 2026 · Artificial Intelligence

Memory-Efficient Algorithms for Large Language Model Inference

The article reviews Coleman Hooper's 2026 Berkeley PhD thesis, which shows that LLM inference is increasingly limited by memory bandwidth and capacity, and proposes a four‑pronged approach—weight quantization, KV‑cache quantization, selective context loading, and multipole attention—to dramatically improve memory efficiency and throughput.

AttentionKV CacheLLM
0 likes · 12 min read
Memory-Efficient Algorithms for Large Language Model Inference
Mingyi World Elasticsearch
Mingyi World Elasticsearch
Aug 7, 2026 · Backend Development

How to Turn Raw Text Data into an Interactive Searchable Dashboard in One Minute for Pre‑sales POCs

The article describes a fully automated pipeline that lets pre‑sales engineers upload a raw CSV/JSON sample, automatically infer mappings, mask sensitive fields, ingest data into Easysearch, generate a searchable, chart‑driven dashboard, and clean up the session with a single click, eliminating the tedious manual preparation that normally dominates POC demos.

Data IngestionEasysearchElasticsearch
0 likes · 14 min read
How to Turn Raw Text Data into an Interactive Searchable Dashboard in One Minute for Pre‑sales POCs
51CTO HarmonyOS Developer Community
51CTO HarmonyOS Developer Community
Aug 7, 2026 · Artificial Intelligence

Building a Medication Plan Workflow for HarmonyOS Intelligent Agents

The article details developing a workflow for adding medication plans in HarmonyOS intelligent agents, covering current time retrieval, LLM-based parameter extraction with a detailed prompt, iterative validation and user prompting for missing info, fixed-option frequency selection, plugin invocation, and error handling for device-dependent plugins.

HarmonyOSLLMMedication Plan
0 likes · 17 min read
Building a Medication Plan Workflow for HarmonyOS Intelligent Agents
Wu Shixiong's Large Model Academy
Wu Shixiong's Large Model Academy
Aug 7, 2026 · Artificial Intelligence

Why a Single -100 Line Determines Who the Multi‑Round SFT Learns to Speak

The article explains how using -100 as an ignored label in PyTorch cross‑entropy loss silently masks non‑assistant tokens, how to locate assistant spans via prefix‑difference, the trade‑offs between supervising only the final reply versus all assistant turns, and the essential pre‑training checks to avoid hidden masking errors in multi‑round SFT.

LLMPyTorchSFT
0 likes · 14 min read
Why a Single -100 Line Determines Who the Multi‑Round SFT Learns to Speak
AI Engineer Programming
AI Engineer Programming
Aug 7, 2026 · Artificial Intelligence

How to Ensure Reliable Structured Outputs in LLM Agents

The article explains why format constraints alone cannot guarantee correct content in LLM agents, compares JSON Mode, Structured Outputs, and Tool Calling, and provides a step‑by‑step engineering guide—including model‑specific quirks, schema validation, retry loops, and layered fallback strategies—to achieve robust structured results.

AgentJSON ModeLLM
0 likes · 13 min read
How to Ensure Reliable Structured Outputs in LLM Agents
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Aug 6, 2026 · Artificial Intelligence

Training‑Free Beats 14B Model: Sonar‑TS Fills Scale Gap in Time‑Series QA

The paper introduces Sonar‑TS, a training‑free neural‑symbolic system that tackles the newly defined NLQ4TSDB problem—natural‑language queries over database‑scale time‑series—by converting shape intents into searchable symbols and verifying candidates with executable code, achieving up to 3.8× higher scores than the strongest Text‑to‑SQL baseline while highlighting remaining challenges in shape understanding.

LLMNatural Language QuerySQL
0 likes · 10 min read
Training‑Free Beats 14B Model: Sonar‑TS Fills Scale Gap in Time‑Series QA
Architect
Architect
Aug 6, 2026 · Artificial Intelligence

Deconstructing TencentDB Agent Memory: How to Keep Agents Accurate Without Being Misled by Errors?

The article analyzes TencentDB Agent Memory’s design, breaking down its three‑stage write‑read‑governance pipeline, four‑level L0‑L3 hierarchy, object types, recall strategies, conflict handling, sharing rules, long‑task traceability, and practical testing guidelines to ensure past information helps future decisions while preserving provenance and error correction.

LLMRecall StrategiesTencentDB
0 likes · 22 min read
Deconstructing TencentDB Agent Memory: How to Keep Agents Accurate Without Being Misled by Errors?
Data Party THU
Data Party THU
Aug 6, 2026 · Artificial Intelligence

What Is an AI Agent Harness and Why It’s Essential Beyond the Model

The article explains how an AI Agent Harness transforms a powerful language model into a reliable, controllable agent by adding tool access, memory, permissions, guardrails, observability, and recovery mechanisms, and outlines its core components, workflow, and a practical customer‑service example.

AIGuardrailsLLM
0 likes · 12 min read
What Is an AI Agent Harness and Why It’s Essential Beyond the Model
TonyBai
TonyBai
Aug 6, 2026 · R&D Management

Rust Says AI Can Review Code but Not Write It: Inside the New LLM Policy

The Rust core teams have published an LLM usage policy that permits AI to assist with reviewing, analyzing, and suggesting code but forbids AI‑generated code creation, outlining strict disclosure rules, higher quality thresholds, and the impact on reviewers, contributors, and issue reporters while comparing approaches taken by Go, Zig, and the Linux kernel.

AILLMPolicy
0 likes · 17 min read
Rust Says AI Can Review Code but Not Write It: Inside the New LLM Policy
Sohu Tech Products
Sohu Tech Products
Aug 5, 2026 · Artificial Intelligence

MemoHarness: The Next Evolution of Agents Happens Outside the Model

MemoHarness proposes an Agent Harness that keeps the language model frozen while learning to adjust external control layers across six editable dimensions, showing measurable gains on terminal, code‑generation, and finance benchmarks but acknowledging limited scale, selective transfer, and cost dependencies.

AI agentsExternal ControlLLM
0 likes · 16 min read
MemoHarness: The Next Evolution of Agents Happens Outside the Model
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Aug 5, 2026 · Artificial Intelligence

Large-Model Memory Panorama: The 3‑D Taxonomy Unveiled by Tsinghua’s Tang Jie Team

This review maps the evolving landscape of large‑model memory, classifying mechanisms along three axes—representation, update dynamics, and persistence—while contrasting implicit and explicit approaches, discussing hybrid designs, and outlining challenges such as write strategies, stability, capacity, and evaluation metrics.

Artificial IntelligenceExplicit MemoryHybrid Models
0 likes · 12 min read
Large-Model Memory Panorama: The 3‑D Taxonomy Unveiled by Tsinghua’s Tang Jie Team
Tencent Technical Engineering
Tencent Technical Engineering
Aug 5, 2026 · Artificial Intelligence

Advanced AI Infra: Making Large Language Models Produce Deterministic Outputs

This article analyzes why LLM inference often yields nondeterministic results, explains how floating‑point addition order, GEMM tiling, Split‑K, RMSNorm, FlashAttention, and NCCL all contribute to batch variance, and details the engineering steps vLLM takes to enforce batch‑invariant execution across GPUs.

Batch InvarianceDeterminismFlashAttention
0 likes · 52 min read
Advanced AI Infra: Making Large Language Models Produce Deterministic Outputs
21CTO
21CTO
Aug 5, 2026 · Industry Insights

Don’t Be Fooled by AI: Why Only 1% of People Truly Win with ChatGPT

The article argues that while ChatGPT and other LLMs appear to democratize expertise, they actually widen the gap between ordinary workers and top specialists, illustrating the point with a programmer’s failure, a fashion designer’s success, and the concept of private‑domain knowledge as the real moat.

AIAutomationLLM
0 likes · 9 min read
Don’t Be Fooled by AI: Why Only 1% of People Truly Win with ChatGPT
FunTester
FunTester
Aug 5, 2026 · Artificial Intelligence

Why AI Alone Won’t Boost Quality: From Speed to Risk Prediction in QA

The article analyzes how AI testing is moving from experimental use to strategic QA governance, emphasizing the need for robust validation processes, multi‑layer verification, risk‑prediction metrics, and collaborative GenAI agents to turn speed gains into genuine quality improvements.

AI testingAutomation GovernanceGenerative AI
0 likes · 10 min read
Why AI Alone Won’t Boost Quality: From Speed to Risk Prediction in QA
21CTO
21CTO
Aug 4, 2026 · Artificial Intelligence

JetBrains Open‑Sources KotlinLLM: LLM‑Driven Smart Macros for Compiled Kotlin

JetBrains has open‑sourced the experimental KotlinLLM IntelliJ IDEA plugin, which introduces LLM‑driven smart macros for Kotlin/JVM projects, addressing runtime delegation latency, external agent workflow complexity, and language integration challenges by providing explicit LLM awareness, source‑level persistence, and zero runtime overhead.

IntelliJ IDEAKotlinLLM
0 likes · 4 min read
JetBrains Open‑Sources KotlinLLM: LLM‑Driven Smart Macros for Compiled Kotlin
Xike
Xike
Aug 4, 2026 · Operations

How We Fixed the AI‑Powered xi‑ops Ops Platform’s Critical Pitfalls

This article walks through the security and reliability pitfalls encountered when integrating large language models into the xi‑ops open‑source operations platform—covering unsafe SQL generation, unauthorized SSH actions, knowledge‑base hallucinations, prompt‑engineered bypasses, and configuration sync issues—and explains the concrete engineering safeguards that were implemented to close each gap.

AI OpsLLMMCP
0 likes · 21 min read
How We Fixed the AI‑Powered xi‑ops Ops Platform’s Critical Pitfalls
DataFunSummit
DataFunSummit
Aug 4, 2026 · Artificial Intelligence

MemoHarness: How Agents Evolve Beyond Model Parameters

MemoHarness expands the notion of self‑evolving agents by keeping the language model frozen while continuously adapting the external control system—context assembly, tool interaction, generation settings, workflow orchestration, memory management, and output validation—demonstrating measurable gains on terminal, code‑generation, and finance tasks, yet highlighting limited scalability and transferability.

AI agentsAgentExperience Learning
0 likes · 17 min read
MemoHarness: How Agents Evolve Beyond Model Parameters
Golang Shines
Golang Shines
Aug 4, 2026 · Information Security

Exploring AI-Assisted Penetration Testing and Vulnerability Discovery

The article analyzes the opportunities and challenges of integrating large language models into penetration testing workflows, presents the design of the AI‑Burp‑Copilot plugin, details its layered architecture, implementation specifics, and real‑world limitations such as LLM nondeterminism and coverage of business‑logic flaws.

AIBurp SuiteLLM
0 likes · 15 min read
Exploring AI-Assisted Penetration Testing and Vulnerability Discovery
PaperAgent
PaperAgent
Aug 4, 2026 · Artificial Intelligence

How Peking University’s Two Papers Redefine Agent Skill Evolution

Two recent Peking University papers, VeriSkill and SESA, demonstrate that treating agent skills as self‑evolving memory—updated from failures via responsibility attribution, lesson abstraction, and failure distillation—yields significant performance gains across verification and search tasks and transfers across models.

AgentLLMProgram Verification
0 likes · 9 min read
How Peking University’s Two Papers Redefine Agent Skill Evolution
JD Cloud Developers
JD Cloud Developers
Aug 4, 2026 · Artificial Intelligence

NaviAgent: Scalable Tool Orchestration for Oxygen Agents via Graph‑Driven Bilevel Planning

The paper introduces NaviAgent, a double‑layer architecture that separates LLM‑based planning from graph‑driven tool navigation, explicitly models API‑parameter dependencies, continuously updates the tool graph with execution feedback, and achieves up to 13.1 % higher task success rates on large‑scale API benchmarks.

AI agentsGraph ModelingLLM
0 likes · 16 min read
NaviAgent: Scalable Tool Orchestration for Oxygen Agents via Graph‑Driven Bilevel Planning
AI Engineer Programming
AI Engineer Programming
Aug 4, 2026 · Artificial Intelligence

Why Agents Call Unneeded Tools and How to Tackle It as a System‑Engineering Problem

The article defines tool hallucination in LLM agents, analyses training bias, context pollution, loop feedback and dialogue inertia as root causes, and proposes multi‑layer defenses—including visibility control, intent verification, runtime gating, architectural isolation, and feedback loops—framed as a system‑engineering challenge rather than mere prompt tweaking.

AgentLLMRuntime Guard
0 likes · 17 min read
Why Agents Call Unneeded Tools and How to Tackle It as a System‑Engineering Problem
Machine Heart
Machine Heart
Aug 3, 2026 · Artificial Intelligence

Does VLA Action Prediction Need an LLM? TurboVLA Achieves 32 Hz with 0.2 B Params on RTX 4090

TurboVLA, a real‑time vision‑language‑action model from Huazhong University of Science and Technology and Huawei, bypasses the large language model bottleneck by directly fusing visual and language features, achieving 32 Hz online action prediction on a single RTX 4090 with only 0.2 B parameters and 0.9 GB VRAM, while maintaining high success rates across LIBERO, RoboTwin 2.0, and real‑robot tasks.

LIBEROLLMRTX 4090
0 likes · 11 min read
Does VLA Action Prediction Need an LLM? TurboVLA Achieves 32 Hz with 0.2 B Params on RTX 4090
SuanNi
SuanNi
Aug 3, 2026 · Artificial Intelligence

DeepSeek V4-Flash Official Release: Open‑Source Model Outperforms V4‑Pro Preview

The DeepSeek V4‑Flash model has been officially released and open‑sourced, delivering performance that surpasses the V4‑Pro preview, rivals Claude Opus‑4.8, ranks second on HuggingFace trends, offers a low price‑per‑token, and tops VulcanBench rankings, while hinting at an upcoming V4‑Pro and AI coding assistant.

AIDeepSeekLLM
0 likes · 3 min read
DeepSeek V4-Flash Official Release: Open‑Source Model Outperforms V4‑Pro Preview
AI Large Model Application Practice
AI Large Model Application Practice
Aug 3, 2026 · Artificial Intelligence

Deep Dive into LLM Wiki Engineering: AI Coding, Obsidian Integration, and RAG Collaboration

This article explains how to build and maintain an LLM‑powered knowledge base (LLM Wiki) for AI coding agents, shows practical workflows using Obsidian and custom agents, and compares the governance‑focused Wiki approach with retrieval‑augmented generation, highlighting trade‑offs, metadata design, and integration patterns.

AI codingAgentLLM
0 likes · 16 min read
Deep Dive into LLM Wiki Engineering: AI Coding, Obsidian Integration, and RAG Collaboration
DeepHub IMBA
DeepHub IMBA
Aug 2, 2026 · Artificial Intelligence

Building a From‑Scratch LLM Training Framework: Full GRPO vs PPO vs DPO Comparison on GSM8K

The article presents a from‑scratch LLM training framework called grpo‑llm, implements GRPO with Trio rollout, FSDP and a C++ reward extension, and conducts a controlled experiment comparing GRPO, PPO and DPO on the GSM8K math‑reasoning benchmark, revealing why DPO outperforms the other two under sparse binary rewards.

DPOGRPOGSM8K
0 likes · 10 min read
Building a From‑Scratch LLM Training Framework: Full GRPO vs PPO vs DPO Comparison on GSM8K
BirdNest Tech Talk
BirdNest Tech Talk
Aug 2, 2026 · Artificial Intelligence

Which LLM Builds the Best Web Gomoku Game? A Comparative Test of the Latest Models

The author prompts six recent large language models to create a web‑based Gomoku game and compares the results across four dimensions—AI opponent, execution method, visual design, and engineering rigor—revealing that Opus 5 and gpt‑5.6‑sol deliver the most complete solutions while deepseek‑v4‑pro and glm‑5.2 excel at quick, single‑file prototypes, and MiniMax‑M3, despite its polished UI, fails at human‑vs‑AI play.

AI comparisonGomokuLLM
0 likes · 13 min read
Which LLM Builds the Best Web Gomoku Game? A Comparative Test of the Latest Models
AI Large-Model Wave and Transformation Guide
AI Large-Model Wave and Transformation Guide
Aug 2, 2026 · Artificial Intelligence

Case Study: Ontology Implementation Pitfalls in a Women’s Apparel Startup

The article critiques a women’s apparel startup’s ontology setup, showing how its reliance on static graphs, SQL queries, and a large language model yields a hard‑coded analysis pipeline rather than true multi‑hop reasoning, and explains why a proper reasoning engine is essential for scenario‑driven inference.

LLMMCPdata inference
0 likes · 6 min read
Case Study: Ontology Implementation Pitfalls in a Women’s Apparel Startup

Federated Dual Ontology: Isolating and Coordinating Two Semantic Domains

The article explains how a federated architecture separates code‑architecture and purchasing constraints into independent domains, injects them into LLM context with per‑domain token budgets, and validates them without merging schemas, demonstrating the approach with concrete directory layouts, configuration code, and experimental results.

ConstraintManagementDomainConfigFederatedGraph
0 likes · 9 min read
Federated Dual Ontology: Isolating and Coordinating Two Semantic Domains
AI Architecture Hub
AI Architecture Hub
Aug 2, 2026 · Artificial Intelligence

Why Stronger Models Need Shorter Prompts: Claude 5 Cuts 80% of System Prompts

Anthropic’s July 2026 release of Claude Opus 5 and Fable 5 demonstrates that trimming more than 80% of system prompts can maintain coding benchmark performance, revealing a shift from bulky prompt engineering to a three‑layer context architecture that assigns minimal, task‑specific information to the model.

AI agentClaude-5Context Engineering
0 likes · 15 min read
Why Stronger Models Need Shorter Prompts: Claude 5 Cuts 80% of System Prompts
DeepHub IMBA
DeepHub IMBA
Aug 1, 2026 · Artificial Intelligence

Estimating the GPU Count Needed to Train a Large Language Model

The article presents a practical scaling‑law based method to estimate the total FLOPs, GPU throughput, and required number of GPUs for training a large language model, showing how to compute these values from model parameters, token count, and target training time, and discusses approximation limits and useful tools.

AICompute EstimationGPU
0 likes · 9 min read
Estimating the GPU Count Needed to Train a Large Language Model
Data Party THU
Data Party THU
Aug 1, 2026 · Artificial Intelligence

10 AI Agent Workflows to Save Teams Hours of Repetitive Work

This article presents ten practical AI agent workflow templates, each with a trigger, context, tools, decision rules, and human checkpoints, showing how to automate tasks like email triage, research briefs, form filling, meeting minutes, support routing, content repurposing, competitor monitoring, invoice reconciliation, CRM updates, and QA review.

AI agentsAutomationLLM
0 likes · 13 min read
10 AI Agent Workflows to Save Teams Hours of Repetitive Work
PaperAgent
PaperAgent
Aug 1, 2026 · Artificial Intelligence

Why LLMs Remember Yet Forget: The Cost of Evolving User Intent

Microsoft Research reveals that large language models excel on static single‑turn tasks but dramatically lose accuracy when user intent evolves across multiple turns, especially during function switches; the study formalizes three intent transition types, proposes a backward‑generation framework, and shows modest gains from memory mechanisms while highlighting the need for active intent recaps.

LLMMemory MechanismUser Intent
0 likes · 12 min read
Why LLMs Remember Yet Forget: The Cost of Evolving User Intent
AI Open-Source Efficiency Guide
AI Open-Source Efficiency Guide
Aug 1, 2026 · Artificial Intelligence

Karpathy’s 10 LLM Coding Rules That Instantly Boost Claude and Codex

The Karpathy‑LLM‑Coding‑Rules repository offers ten executable, bilingual rules that constrain AI coding agents like Claude and Codex, providing clear validation criteria and anti‑pattern names to prevent over‑refactoring, hidden bugs, and unbounded dependencies, and can be dropped into a project with a single file copy.

AI agentsClaudeCodex
0 likes · 10 min read
Karpathy’s 10 LLM Coding Rules That Instantly Boost Claude and Codex
Linyb Geek Road
Linyb Geek Road
Aug 1, 2026 · Artificial Intelligence

Practical Guide to Cutting LLM Token Costs

This article systematically explains how large‑language‑model token pricing works, identifies eight high‑consumption usage patterns, presents nine actionable optimization principles, and offers a tiered model‑selection framework so engineering teams can reduce token spend by up to 80% without sacrificing result quality.

Batch ProcessingCachingLLM
0 likes · 22 min read
Practical Guide to Cutting LLM Token Costs
AI Engineering
AI Engineering
Jul 31, 2026 · Artificial Intelligence

Extract JSON from PDFs with Natural Language Using Unstract – No More Regex

Unstract is an open‑source platform that lets you describe desired fields in plain language and leverages LLMs to turn invoices, bank statements, KYC forms and other unstructured PDFs into clean JSON ready for database ingestion, eliminating the need for regex or custom templates.

JSON outputLLMPDF extraction
0 likes · 4 min read
Extract JSON from PDFs with Natural Language Using Unstract – No More Regex
AntData
AntData
Jul 31, 2026 · Artificial Intelligence

When Expert Experience Can Be Quantified: How Rubrics Become Data Assets for LLM Inference Training

The article analyzes how combining formal verification with expert‑derived Rubrics provides fine‑grained process supervision for large language models, presents the CRAFT data‑production pipeline, and shows experimental gains on math and medical benchmarks using Rubric‑driven RL, SFT, and alternating RL‑SFT training.

Formal VerificationLLMRubrics
0 likes · 23 min read
When Expert Experience Can Be Quantified: How Rubrics Become Data Assets for LLM Inference Training
Java Tech Workshop
Java Tech Workshop
Jul 31, 2026 · Backend Development

Building a Pluggable LLM Gateway with Spring Boot

The article explains why enterprise applications need a unified large‑model gateway, outlines a layered architecture using Spring Boot 3.x, strategy pattern and interface‑driven design, and provides production‑grade features such as retry, circuit‑breaker, token accounting and easy addition of new model providers.

LLMOpenFeignResilience4j
0 likes · 19 min read
Building a Pluggable LLM Gateway with Spring Boot
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 30, 2026 · Artificial Intelligence

From LLM Rollout to Agentic Rollout: Design Insights and Lessons for an Agentic RL Training Framework

The article analyzes the transition from single‑step LLM rollouts to multi‑step Agentic RL rollouts, compares coupled and decoupled architectures, details the roles of Controller, Runtime Manager, Gateway and LLM Server, and discusses token‑level consistency, trajectory reconstruction, and scalability strategies for a production‑grade training pipeline.

Agentic RLLLMRollout Architecture
0 likes · 25 min read
From LLM Rollout to Agentic Rollout: Design Insights and Lessons for an Agentic RL Training Framework
ThinkingAgent
ThinkingAgent
Jul 29, 2026 · Artificial Intelligence

How Tokenizers and Embeddings Encode Language – The Mechanics Behind LLMs

The article explains how different tokenization strategies (word, character, subword) affect token counts, model cost, context length, and multilingual fairness, and details the engineering trade‑offs of BPE, WordPiece, Unigram, SentencePiece, special tokens, and embedding matrices in large language models.

AIEmbeddingLLM
0 likes · 23 min read
How Tokenizers and Embeddings Encode Language – The Mechanics Behind LLMs
Linyb Geek Road
Linyb Geek Road
Jul 29, 2026 · Artificial Intelligence

Why Adding More Documents Can Degrade RAG Answers

The article explains that stuffing a RAG system with many overlapping or conflicting documents consumes tokens, slows responses, and introduces noise that prevents the model from correctly using the most relevant evidence, ultimately worsening answer quality.

Evidence RankingLLMRAG
0 likes · 14 min read
Why Adding More Documents Can Degrade RAG Answers
DataFunTalk
DataFunTalk
Jul 28, 2026 · Artificial Intelligence

What Is an Agent Harness? A Deep Dive into AI Agent Architecture

The article explains that an Agent Harness is the full software infrastructure surrounding a large language model—handling orchestration loops, tool integration, memory, context management, error handling, and security—and shows how production‑grade harnesses, defined by Anthropic, OpenAI and LangChain, consist of twelve components, with detailed design trade‑offs and practical examples.

AI agentsLLMagent harness
0 likes · 21 min read
What Is an Agent Harness? A Deep Dive into AI Agent Architecture
PaperAgent
PaperAgent
Jul 28, 2026 · Artificial Intelligence

Inside Anthropic’s New Graph Engineering Methodology for Multi‑Agent Systems

Anthropic’s recent 12‑page playbook and 2‑hour workshop detail a Graph Engineering pipeline that replaces costly context‑window communication with a shared knowledge graph, covering why windows fail, a four‑stage Claude API workflow, extraction rules, entity resolution, graph assembly, multi‑hop querying, integration into five agent modes, cost analysis, scaling strategies, and guidance on when not to use a knowledge graph.

Agentic AIAnthropicClaude API
0 likes · 14 min read
Inside Anthropic’s New Graph Engineering Methodology for Multi‑Agent Systems
Ray's Galactic Tech
Ray's Galactic Tech
Jul 27, 2026 · Artificial Intelligence

From Zero to One: Building an AI Requirement Analysis Engine – Full Technical Walkthrough of the Skills Architecture

This article details how to engineer a production‑grade AI requirement‑analysis engine using a modular Skills framework, covering problem definition, architecture, compilation pipeline, conflict detection, impact analysis, scalability, multi‑tenant governance, and deployment best practices.

AIKnowledge RetrievalLLM
0 likes · 34 min read
From Zero to One: Building an AI Requirement Analysis Engine – Full Technical Walkthrough of the Skills Architecture
BanTech Think Tank
BanTech Think Tank
Jul 27, 2026 · Operations

How LLM‑Powered Intelligent Workflow Orchestration Accelerates Financial Digital Transformation

The article analyzes the shortcomings of traditional rule‑based banking workflow orchestration, proposes an LLM‑and‑RAG‑driven end‑to‑end framework that understands requirements, generates standardized process configurations, and validates them automatically, and reports a 50% boost in development efficiency and 99.99% stability in production.

Financial TechnologyLLMProcess Automation
0 likes · 14 min read
How LLM‑Powered Intelligent Workflow Orchestration Accelerates Financial Digital Transformation
DataFunTalk
DataFunTalk
Jul 27, 2026 · Artificial Intelligence

MemoHarness: The Next Evolution of Agents Happens Outside the Model

MemoHarness introduces an Agent Harness that keeps the language model frozen while iteratively optimizing the surrounding control system across six dimensions, showing measurable gains on terminal automation, code generation, and financial analysis tasks, yet acknowledges limited experimental scale and selective transferability.

AI agentsExternal ControlLLM
0 likes · 16 min read
MemoHarness: The Next Evolution of Agents Happens Outside the Model
ThinkingAgent
ThinkingAgent
Jul 27, 2026 · Artificial Intelligence

The Awakening of Large Models: From Classic Language Modeling to Generative AI

This article traces the 56‑year evolution of language models—from ELIZA’s rule‑based scripts and N‑gram statistics to neural embeddings, RNNs, Transformers and the seven‑layer ChatGPT architecture—explaining why the simple next‑token probability definition has remained the core of generative AI, how autoregressive factorization drives training, generation and decoding, why hallucinations arise, and what engineering trade‑offs matter in production.

ChatGPTLLMLarge Language Models
0 likes · 26 min read
The Awakening of Large Models: From Classic Language Modeling to Generative AI
AI Engineer Programming
AI Engineer Programming
Jul 26, 2026 · Artificial Intelligence

Agent Development Lifecycle (ADLC): Vendor‑Neutral Guide to Build, Test, Deploy, Monitor, and Govern AI Agents

This note outlines a vendor‑agnostic Agent Development Lifecycle (ADLC) that extends traditional SDLC with five stages—Build, Test, Deploy, Monitor, and Govern—detailing layer‑wise tooling choices, evaluation strategies, deployment infrastructure, observability practices, and governance concerns for modern AI agents.

AI lifecycleAgentOpsLLM
0 likes · 15 min read
Agent Development Lifecycle (ADLC): Vendor‑Neutral Guide to Build, Test, Deploy, Monitor, and Govern AI Agents
The Dominant Programmer
The Dominant Programmer
Jul 26, 2026 · Artificial Intelligence

Building Smart Agents with Spring AI Alibaba: A Hands‑On Guide

This article walks through the Spring AI Alibaba Agent Framework (v1.1.2.0), explaining the ReAct reasoning‑acting loop, core APIs, configuration, code examples, testing commands, and common troubleshooting steps so developers can quickly create LLM‑driven agents with tool‑calling and memory support.

Agent FrameworkAlibabaJava
0 likes · 15 min read
Building Smart Agents with Spring AI Alibaba: A Hands‑On Guide
Ubuntu
Ubuntu
Jul 26, 2026 · Artificial Intelligence

Why Leading AI Coding Agents Like Claude Code, Pi, and OpenCode Are Built with JavaScript/TypeScript

Despite Python’s dominance in AI, the top AI coding agents converge on TypeScript + Node.js because five concrete engineering trade‑offs—event‑loop alignment with ReAct, native streaming support, a rich npm ecosystem, TypeScript being the LLM’s “native language”, and hot‑pluggable dynamic imports—make JavaScript the optimal stack, with clear exceptions for certain workloads.

AI agentsJavaScriptLLM
0 likes · 12 min read
Why Leading AI Coding Agents Like Claude Code, Pi, and OpenCode Are Built with JavaScript/TypeScript
DataFunTalk
DataFunTalk
Jul 26, 2026 · Artificial Intelligence

Agent Harness Deep Dive: Unpacking the Architecture Behind AI Agents

The article dissects the concept of an Agent Harness, distinguishes it from the agent itself, outlines three engineering layers, enumerates twelve production‑grade components, walks through a full execution loop, and compares how major frameworks implement these ideas.

AI agentsLLMReact
0 likes · 20 min read
Agent Harness Deep Dive: Unpacking the Architecture Behind AI Agents
AI Engineer Programming
AI Engineer Programming
Jul 26, 2026 · Artificial Intelligence

Analyzing the grill‑me Agent Skills Repository: Making Probabilistic LLMs Deterministic

The article dissects Matt Pocock’s skills repository, explaining how a set of atomic, editable, composable Agent Skills—driven by structured grilling, shared vocabularies, TDD loops, and design checkpoints—turns the inherently probabilistic nature of LLM‑based programming into a repeatable, deterministic workflow while highlighting practical limits and best‑practice patterns.

AgentAutomationLLM
0 likes · 22 min read
Analyzing the grill‑me Agent Skills Repository: Making Probabilistic LLMs Deterministic
PaperAgent
PaperAgent
Jul 25, 2026 · Artificial Intelligence

Inside Claude Code and Codex: Dissecting the Six Core Components of a Coding Agent

The article breaks down the architecture of coding agents like Claude Code and Codex into six essential components—Live Repo Context, Prompt Cache, Tools, Context Management, Session Memory, and Bounded Subagents—explaining how each layer of the Agent Harness transforms similar LLMs into markedly different, more capable systems.

LLMPrompt CachingSession Memory
0 likes · 12 min read
Inside Claude Code and Codex: Dissecting the Six Core Components of a Coding Agent
Machine Heart
Machine Heart
Jul 25, 2026 · Artificial Intelligence

Eight LLM Phone Agents Commit Real‑World Fraud on Devices – New Security Dataset

The researchers integrated eight LLM‑based phone agents into real smartphones, evaluated them across 31 popular apps using the newly created BadPhoneAgent dataset, and found alarmingly low safety awareness yet high success rates and human‑level speed in executing malicious tasks such as fraud and illicit purchases.

AI safetyLLMdataset
0 likes · 8 min read
Eight LLM Phone Agents Commit Real‑World Fraud on Devices – New Security Dataset
Weekly Large Model Application
Weekly Large Model Application
Jul 25, 2026 · Artificial Intelligence

How Pipeline Parallelism Cuts AI Voice Latency Below 700 ms

The article explains that keeping end‑to‑end voice‑assistant latency under 700 ms requires a co‑designed pipeline—streaming STT, speculative LLM, and streaming TTS—rather than faster individual models, and it details concrete component choices, budget allocations, and common pitfalls.

AI voiceLLMSTT
0 likes · 9 min read
How Pipeline Parallelism Cuts AI Voice Latency Below 700 ms
Yunqi AI+
Yunqi AI+
Jul 24, 2026 · Artificial Intelligence

Designing a Production-Ready Business Analysis Skill for Enterprise AI

The article outlines a deterministic, modular architecture for enterprise AI business‑analysis agents, detailing how to split reports into audited modules, assign clear tool contracts, perform rigorous attribution, generate evidence‑backed insights, and implement robust review, versioning, and evaluation practices.

AIAgentAttribution
0 likes · 21 min read
Designing a Production-Ready Business Analysis Skill for Enterprise AI
AI Engineering
AI Engineering
Jul 24, 2026 · Artificial Intelligence

Andrew Ng’s OpenWorker: An Out‑of‑the‑Box AI Agent Built for Getting Real Work Done

OpenWorker, the newly open‑sourced AI agent announced by Andrew Ng, lets users specify desired outcomes and automatically breaks tasks into steps, invokes selected LLMs and tools, and delivers completed results—supporting 25+ integrations, local data handling, model‑agnostic operation, and a safety‑first approval workflow.

AI agentAndrew NgAutomation
0 likes · 4 min read
Andrew Ng’s OpenWorker: An Out‑of‑the‑Box AI Agent Built for Getting Real Work Done
AI Programming Lab
AI Programming Lab
Jul 23, 2026 · Artificial Intelligence

How Codex and Claude Code Compress Context: Mechanisms, Experiments, and Performance

The article analyzes Codex's opaque, encrypted compaction items versus Claude Code's transparent summaries, explains trigger mechanisms, details a reverse‑engineering prompt‑injection experiment, and presents a benchmark where native server compression achieves 100% accuracy while plain text summaries lag behind.

AnthropicClaude CodeCodex
0 likes · 11 min read
How Codex and Claude Code Compress Context: Mechanisms, Experiments, and Performance
DataFunTalk
DataFunTalk
Jul 23, 2026 · Artificial Intelligence

Deep Dive into Agent Harness: Dissecting the Architecture Behind AI Agents

The article explains that an Agent Harness is the full software infrastructure surrounding a large language model—handling orchestration loops, tool integration, memory, context management, state persistence, error handling, safety guards, and validation—showing why harness design, not model size, determines production‑grade agent performance.

AI agentsContext EngineeringLLM
0 likes · 19 min read
Deep Dive into Agent Harness: Dissecting the Architecture Behind AI Agents
DeWu Technology
DeWu Technology
Jul 23, 2026 · Artificial Intelligence

When Engineers Cross Boundaries: How “Boundary‑Breaking” Boosted Problem Solving at Dewu

The Dewu tech team’s recent “boundary‑crossing” incidents—engineers skipping formal specs to talk directly with users and operations—led to deeper problem understanding, rapid prototyping, and measurable improvements such as 60% automated answers, 30% support load reduction, and 50% faster onboarding.

AI assistantAI platformLLM
0 likes · 7 min read
When Engineers Cross Boundaries: How “Boundary‑Breaking” Boosted Problem Solving at Dewu
Wu Shixiong's Large Model Academy
Wu Shixiong's Large Model Academy
Jul 23, 2026 · Artificial Intelligence

How to Prevent RAG from Hallucinating When No Answer Exists – Beyond Simple Similarity Thresholds

The article explains why a plain similarity‑threshold check cannot reliably stop Retrieval‑Augmented Generation from fabricating answers, introduces a four‑stage evidence‑control framework, details how to calibrate thresholds with balanced positive and negative samples, and outlines concrete actions for handling insufficient evidence.

LLMRAGanswerability
0 likes · 21 min read
How to Prevent RAG from Hallucinating When No Answer Exists – Beyond Simple Similarity Thresholds
Machine Heart
Machine Heart
Jul 22, 2026 · Artificial Intelligence

Google Unveils Three New Gemini Flash Models as Gemini 3.5 Pro Remains Delayed

Google introduced Gemini 3.6 Flash, Gemini 3.5 Flash‑Lite, and Gemini 3.5 Flash Cyber, detailing their efficiency gains, benchmark improvements, lower pricing, and limited release strategies while noting that Gemini 3.5 Pro is still postponed and Gemini 4 is already in training.

Flash modelsGeminiGoogle AI
0 likes · 9 min read
Google Unveils Three New Gemini Flash Models as Gemini 3.5 Pro Remains Delayed
Java Tech Enthusiast
Java Tech Enthusiast
Jul 22, 2026 · Artificial Intelligence

When New LLMs Impress, Their Flaws Quickly Disappoint

The author tests CodeX and GPT5.6‑Sol on a multi‑task directory workflow and finds simple yet puzzling errors, then observes Fable5 failing on basic CSS tweaks, linking both issues to catastrophic forgetting and hallucination in large language models.

CodexFable5GPT-5.6
0 likes · 7 min read
When New LLMs Impress, Their Flaws Quickly Disappoint
PaperAgent
PaperAgent
Jul 22, 2026 · Artificial Intelligence

Inside GPT‑5.6’s Dropdown: How Six Leading LLMs Tune Their Reasoning Effort

The article dissects Sebastian Raschka’s “Controlling Reasoning Effort in LLMs”, explains GPT‑5.6’s multi‑level effort settings, clarifies the notion of reasoning models, outlines training vs. inference scaling, details RLVR recipes, and compares the post‑training formulas of six open‑source flagship LLMs.

GPT-5.6LLMOpen‑source Models
0 likes · 12 min read
Inside GPT‑5.6’s Dropdown: How Six Leading LLMs Tune Their Reasoning Effort
Wu Shixiong's Large Model Academy
Wu Shixiong's Large Model Academy
Jul 22, 2026 · Artificial Intelligence

How to Handle Long Conversation History: Beyond Full Prompt or Recent Rounds

The article explains that effective conversation memory for LLMs requires classifying information into static knowledge, short‑term context, and long‑term memory, defining a full lifecycle for each entry, and implementing strict storage, retrieval, update, and deletion policies rather than simply concatenating all history or keeping only the latest turns.

LLMRAGconversation memory
0 likes · 24 min read
How to Handle Long Conversation History: Beyond Full Prompt or Recent Rounds
AI Engineer Programming
AI Engineer Programming
Jul 22, 2026 · Artificial Intelligence

Is Prompt Engineering Dead? A Deep Dive into Harness, Context Assembly, and Token Generation

The article examines why traditional prompt engineering is no longer sufficient in production AI systems, detailing how harness layers, context reassembly, tool orchestration, token generation methods, training objectives, and architecture choices transform a simple prompt into a complex, multi‑stage workflow that demands robust, system‑level design.

HarnessLLMMulti-Token Prediction
0 likes · 16 min read
Is Prompt Engineering Dead? A Deep Dive into Harness, Context Assembly, and Token Generation
TechVision Expert Circle
TechVision Expert Circle
Jul 21, 2026 · Artificial Intelligence

Apple’s New Siri Public Beta Redefines Mobile Assistants with LLM‑Based Agent Architecture

Apple’s July 2026 public beta of Siri replaces its legacy intent‑based pipeline with a large‑language‑model‑driven agent architecture, introducing multimodal perception, persistent memory, and a three‑tier edge‑cloud inference system that reshapes mobile assistants while emphasizing privacy through on‑device processing and differential‑privacy techniques.

Agent ArchitectureAppleEdge AI
0 likes · 13 min read
Apple’s New Siri Public Beta Redefines Mobile Assistants with LLM‑Based Agent Architecture