PaperAgent
Author

PaperAgent

Daily updates, analyzing cutting-edge AI research papers

302
Articles
1
Likes
1.8k
Views
0
Comments
Recent Articles

Latest from PaperAgent

100 recent articles max
PaperAgent
PaperAgent
Aug 14, 2026 · Artificial Intelligence

DeepSeek Harness Open-Source Review: Surprising Insights Beyond Its Plugin System

The article provides a detailed technical walkthrough of DeepSeek Harness (dsh), highlighting its plugin‑centric architecture, full‑trajectory visibility, performance metrics, installation steps, four UI modes, the underlying Cordis framework, the capability‑seam design, and community‑contributed plugins, all illustrated with concrete examples and code.

AI AgentsCordis frameworkDeepSeek Harness
0 likes · 7 min read
DeepSeek Harness Open-Source Review: Surprising Insights Beyond Its Plugin System
PaperAgent
PaperAgent
Aug 13, 2026 · Artificial Intelligence

First Community Benchmarks of DeepSeek V4 Pro, Qwen 3.8 Max, and Grok 4.6

The community quickly tested three newly released LLMs—DeepSeek V4 Pro, Qwen 3.8 Max, and Grok 4.6—across 3D scene generation, Flappy game creation, and airplane‑animation tasks, comparing quality, speed, and cost to reveal each model’s strengths and trade‑offs.

AIDeepSeekGrok
0 likes · 5 min read
First Community Benchmarks of DeepSeek V4 Pro, Qwen 3.8 Max, and Grok 4.6
PaperAgent
PaperAgent
Aug 11, 2026 · Artificial Intelligence

Long-Horizon Tasks Jump 98%: Introducing Stanford’s Skill‑Native LLM

Researchers introduce Skill‑Entropy, a metric quantifying the difficulty of switching between reasoning skills in long‑horizon tasks, build the 558‑skill Skill²‑Bench, and show that Skill‑Entropy‑RL training dramatically improves cross‑skill performance of LLMs such as Qwen3, closing the gap observed in standard benchmarks.

Cross‑Skill ReasoningLLM BenchmarkingQwen3
0 likes · 12 min read
Long-Horizon Tasks Jump 98%: Introducing Stanford’s Skill‑Native LLM
PaperAgent
PaperAgent
Aug 9, 2026 · Artificial Intelligence

How to Build a Fully Local Coding Agent: Best Practices and Benchmarks

This tutorial walks through assembling a completely offline coding agent using open‑source tools and open‑weight models, evaluates Qwen‑Code versus Codex and Claude Code harnesses with speed, capability and token‑usage benchmarks, and provides security‑audit and configuration guidance.

HarnessOllamaQwen3.6
0 likes · 13 min read
How to Build a Fully Local Coding Agent: Best Practices and Benchmarks
PaperAgent
PaperAgent
Aug 9, 2026 · Artificial Intelligence

Tsinghua Unveils Two Breakthrough Papers on LLM Agent Skills

The article reviews Tsinghua University's two new papers—GSE, which introduces a global skill‑relation graph, clustering, and replay verification to make agent skills continuously improve, and SkillSentry, which uses ability contracts and adaptive honey‑world testing to ensure skill safety—detailing their methods, experimental results, and practical implications.

AI safetyAgentGSE
0 likes · 8 min read
Tsinghua Unveils Two Breakthrough Papers on LLM Agent Skills
PaperAgent
PaperAgent
Aug 7, 2026 · Artificial Intelligence

Jeff Dean Launches Discovery Loop to Automate the AI Experimental Cycle

Jeff Dean, together with Sanjay Ghemawat, Oriol Vinyals, and Quoc Le, founded Discovery Loop, a public‑benefit corporation that seeks to replace manual AI prompting with an automated experimental loop—proposing, running, and learning from thousands of ML experiments to accelerate discovery across scientific domains.

AI AgentsAI automationDiscovery Loop
0 likes · 7 min read
Jeff Dean Launches Discovery Loop to Automate the AI Experimental Cycle
PaperAgent
PaperAgent
Aug 7, 2026 · Artificial Intelligence

OpenMLE: Tsinghua’s Self‑Evolving MLE System Pushes 35B Model Past GPT‑5.5

The article introduces OpenMLE, an open‑source full‑stack system for recursive self‑improvement (RSI) research, showing how a 35B Frontis‑MA1 model improves its Medal Average from 39.39% to 71.21% on MLE‑Bench Lite, surpasses GPT‑5.5+Codex, and details the mechanism hierarchy, task‑curation gym, trainable evolution operators, and experimental evidence that training and search gains combine additively.

Evolutionary SearchFrontis-MA1MLE-Bench Lite
0 likes · 19 min read
OpenMLE: Tsinghua’s Self‑Evolving MLE System Pushes 35B Model Past GPT‑5.5
PaperAgent
PaperAgent
Aug 6, 2026 · Artificial Intelligence

What Research Directions Are Worth Pursuing After Reviewing 407 Large Model Papers?

The author curates a collection of 407 recent large‑model papers—264 frontier works across six innovation paths and 143 top‑conference papers—classifies them into 14 hot sub‑topics, and explains how labs can match these directions to their available compute, data, and time resources.

AI researchLarge ModelsRetrieval-Augmented Generation
0 likes · 4 min read
What Research Directions Are Worth Pursuing After Reviewing 407 Large Model Papers?
PaperAgent
PaperAgent
Aug 4, 2026 · Artificial Intelligence

How Peking University’s Two Papers Redefine Agent Skill Evolution

Two recent Peking University papers, VeriSkill and SESA, demonstrate that treating agent skills as self‑evolving memory—updated from failures via responsibility attribution, lesson abstraction, and failure distillation—yields significant performance gains across verification and search tasks and transfers across models.

AgentLLMProgram Verification
0 likes · 9 min read
How Peking University’s Two Papers Redefine Agent Skill Evolution