Tagged articles

Model Editing

6 articles · Page 1 of 1
Geek Labs
Geek Labs
Sep 9, 2026 · Artificial Intelligence

Heretic: Automated Model Surgery Removes Refusal Without Damaging Capabilities

This article explains how Heretic automates directional ablation to permanently remove refusal behavior from language models by editing model weights, contrasting it with jailbreaks and steering vectors, and details its use of Optuna optimization with KL divergence to preserve model capabilities, situating it within a broader taxonomy of model modification techniques.

AbliterationHereticKL Divergence
0 likes · 18 min read
Heretic: Automated Model Surgery Removes Refusal Without Damaging Capabilities
ThinkingAgent
ThinkingAgent
Aug 26, 2026 · Artificial Intelligence

Trustworthy AI: Security, Explainability, Governance, and the 2026 Technical Ceiling

The article examines how increasingly capable large language models transition from merely avoiding harmful output to preventing harmful actions, outlining threat modeling, prompt injection, data‑pipeline attacks, architectural controls, interpretability, governance frameworks, and the unresolved technical limits that persist through 2026.

AI SafetyGovernanceModel Editing
0 likes · 38 min read
Trustworthy AI: Security, Explainability, Governance, and the 2026 Technical Ceiling
Network Intelligence Research Center (NIRC)
Network Intelligence Research Center (NIRC)
Nov 11, 2025 · Artificial Intelligence

What Is Mechanistic Interpretability and Why It Matters for Large Language Models

The article defines mechanistic interpretability as reverse‑engineering LLMs to reveal how they represent knowledge and make decisions, explains its importance for transparency, risk mitigation, and model improvement, and surveys key techniques such as causal tracing, zero‑making, noise‑making, and logit‑lens methods with illustrative examples.

Large Language ModelsMechanistic InterpretabilityModel Editing
0 likes · 8 min read
What Is Mechanistic Interpretability and Why It Matters for Large Language Models
DataFunTalk
DataFunTalk
Jul 4, 2025 · Artificial Intelligence

How to Edit Large Language Models: Techniques, Metrics, and Challenges

This article explains model editing—injecting or updating knowledge in AI models—distinguishes it from post‑training, outlines reliability, generalization and locality metrics, and surveys both parameter‑free (e.g., IKE) and parameter‑based methods such as ROME, hypernetworks, and MEND, highlighting practical challenges.

MENDModel EditingRome
0 likes · 10 min read
How to Edit Large Language Models: Techniques, Metrics, and Challenges
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Aug 20, 2024 · Artificial Intelligence

How DAFNet Enables Efficient Sequential Editing of Large Language Models

This article introduces DAFNet, a dynamic auxiliary fusion framework that enables efficient sequential editing of large language models by injecting knowledge with reduced resource costs while preserving model reliability, generalization, and mitigating hallucination, and details its dataset, architecture, and evaluation results.

AI researchModel Editingdynamic auxiliary fusion
0 likes · 10 min read
How DAFNet Enables Efficient Sequential Editing of Large Language Models