SingProbe: Endogenous LLM Guardrails Cut Overhead to 0.5% While Matching External Baselines

SingProbe introduces endogenous runtime guardrails that integrate risk detection directly into LLM inference by reusing hidden states, achieving comparable safety and hallucination detection to external guardrails while adding less than 0.5% decode overhead across 29 open-source models.

AntTech
AntTech
AntTech
SingProbe: Endogenous LLM Guardrails Cut Overhead to 0.5% While Matching External Baselines

When medical LLMs enter real-world scenarios like diagnosis assistance and medication consulting, a natural question arises: can we trust the model's answers? Users worry whether the model correctly understands high-risk medical queries, gives unsafe advice, or fabricates guidelines and dosages. To address this, systems need continuous judgment during generation: whether user intent carries risk, whether responses are safe, and whether factual grounding is reliable — traditionally incurring heavy overhead.

SingProbe is an endogenous runtime guardrail for LLMs. It moves guardrails from independent audit modules into the model's inference process, producing internal safety signals synchronously with token generation while reducing overhead to a fraction of traditional approaches. Traditional external guardrails are like "model finishes, then a separate reviewer re-listens"; SingProbe is like "a built-in polygraph emits risk signals as the model speaks." The former requires extra model forward passes, VRAM, and service orchestration, typically adding >20% latency; the latter reuses the same inference's intermediate states, avoids re-encoding text, and keeps decode overhead <0.5%. Endogenous guardrails can work standalone or complement external ones.

Three direct impacts:

Different form: Risk judgment sits inside the LLM inference chain, not as an external audit model.

Lighter speed: Production tests on Ling-3.0-flash show <0.5% extra decode overhead.

Usable capability: On user intent, response safety, streaming detection, and hallucination detection, results match or exceed corresponding public external guardrails.

The current release adapts to 29 mainstream open-source models including Ling-3.0 series, GLM-5.2/5.3, Qwen, DeepSeekV4 , and integrates into SGLang and vLLM inference pipelines. For systems already using these frameworks, SingProbe adds risk logits alongside token logits rather than requiring a separate guardrail service.

Why Guardrails Need to Be Endogenous

Existing guardrails typically use external deployment: main model generates, separate guardrail model audits. This mature, general approach faces three unavoidable problems in production:

System cost from independent deployment: External guardrails re-process user input or model output, adding a full semantic encoding and judgment pass beyond the main model. This brings extra parameters, VRAM usage, service orchestration, and cross-service communication/sync overhead.

Capability mismatch between guardrail and main model: Many guardrail models are smaller than the monitored LLM. Facing long contexts, complex reasoning, or multi-step agent tasks, they may fail to stably understand what the main model generates. Keeping external guardrails pace with the main model's capability, context length, and task complexity is itself challenging.

Safety signal latency: Traditional guardrails often judge safety after full response generation; even with streaming detection, to amortize external guardrail inference cost they often check by text chunks rather than per token. This makes timely intervention at the moment risk appears difficult.

Recent industry practice quantifies these costs:

OpenAI (Aug 2026) disclosed a multi-stage chain-of-thought monitoring system for high-risk workloads: activating classifiers per sampled token, escalating suspicious behavior to higher-compute investigation systems. Estimated compute overhead ~20% of monitored inference compute, varying widely across training/evaluation loads [1].

NVIDIA NeMo Guardrails (2025) reported: adding content moderation, jailbreak detection, and topic control increased average latency from 0.91s to 1.44s, single-interaction throughput dropped from 112.9 to 98.7 tokens/s [2].

A 2026 systematic study of 13 guardrails found most add up to ~0.5s extra latency; reasoning-dependent guardrails incur higher overhead [3].

These systems confirm a reality: external guardrails aren't just "add one more review." They bring extra inference cost, affect signal timing, and as main model tasks grow more complex, whether the guardrail itself has sufficient understanding becomes a deployment challenge.

Meanwhile, research shows LLM hidden states during inference already encode rich safety-discriminative information — internal representations implicitly record risk features when generating safe/unsafe content. This provides theoretical basis for endogenous guardrails: since risk signals exist in the main model's internal states, systems don't need extra models to re-understand text; lightweight probes can directly "read out" these signals. SingProbe chooses endogenous guardrails so guardrails become side-signals produced synchronously with token logits during autoregressive decoding, not "secondary audits" outside the main model.

Architecture Comparison

The following comparison contrasts traditional external guardrails with SingProbe's endogenous approach:

Risk signal source: Traditional — re-encode and reason over text again; SingProbe — produced synchronously during decoding.

Integration with generation: Traditional — independent model or service chain; SingProbe — plugged into autoregressive decoding as side-signal.

Detection granularity: Traditional — typically whole segment or text chunks; SingProbe — per-token causal update.

Current implementation: Traditional — extra model forward pass; SingProbe — reuse hidden states + lightweight prediction heads.

Extra compute: Traditional — depends on external guardrail model & call style, usually >20%; SingProbe — decode overhead <0.5%, nearly negligible.

By reusing hidden states already produced during main model inference and attaching lightweight prediction heads, risk scores update alongside autoregressive decoding without extra text encoding cost. For streaming systems, this means safety detection no longer forces a choice between "finer detection frequency" and "lower running cost."

Each SingProbe model outputs multiple scores per token position: user intent category, response unsafety score, hallucination risk score — supporting three tasks in one interface:

User intent recognition: Identify risky vs safe intents, signaling input-side policies.

Response safety monitoring: Track risk in generated responses, locate when risk begins.

Hallucination risk detection: Continuously estimate factual reliability risk of generated content.

Since each score depends only on the already-generated prefix, SingProbe's judgments are immediate: systems need not wait for full response completion to feed risk signals to downstream policies like alerting, termination, retry, or constrained decoding.

Does Endogenous Guarding Stay Accurate?

Lightweight shouldn't mean sacrificing judgment. Systematic evaluation across user intent classification, response safety classification, streaming safety detection, and hallucination detection. Metrics: F1, AUC, R-AUC, T-AUC (all higher-is-better). Reference baselines are strongest/most representative public baselines per task from technical reports.

Evaluation Results

User Intent Classification (6 benchmarks, F1): Ling-3.0-tiny-singprobe 0.8561; Ling-3.0-flash-singprobe 0.8674; Reference baseline YuFeng-XGuard-Reason-8B: 0.8714 .

Response Safety Classification (8 benchmarks, F1): Ling-3.0-tiny-singprobe 0.8508; Ling-3.0-flash-singprobe 0.8728 ; Reference baseline Qwen3Guard-Gen-8B-strict: 0.8604.

Streaming Safety Detection (3 benchmarks, R-AUC): Ling-3.0-tiny-singprobe 0.9888 ; Ling-3.0-flash-singprobe 0.9887; Reference baseline Qwen3Guard-Stream-8B: 0.9640.

Streaming Safety Detection (3 benchmarks, T-AUC): Ling-3.0-tiny-singprobe 0.9479; Ling-3.0-flash-singprobe 0.9481 ; Reference baseline Qwen3Guard-Stream-8B: 0.8893.

Hallucination Detection (6 benchmarks, AUC): Ling-3.0-tiny-singprobe 0.7765; Ling-3.0-flash-singprobe 0.8012 ; Reference baseline DRIFT: 0.8000.

Four conclusions:

On response safety classification, Ling-3.0-flash-singprobe average exceeds Qwen3Guard-Gen-8B-strict baseline.

On streaming safety detection, SingProbe more stably distinguishes "still-safe prefixes" from "truly risky continuations."

On hallucination detection, flash version matches DRIFT baseline.

Beyond these detection capabilities, decode-phase overhead remains <0.5% .

This addresses the common capability-mismatch problem of external guardrails: guardrails shouldn't be fixed-capability external auditors but should scale with the main model's context length, reasoning ability, and task complexity. SingProbe trains separate probes for Ling-3.0-tiny and Ling-3.0-flash; experiments show as base model scales from tiny to flash, response safety classification and hallucination detection improve synchronously. Endogenous guardrails can gain stronger risk perception via main model scaling, rather than maintaining an increasingly hard-to-match external audit model.

Can Risks Be Caught at the Exact Moment They Appear?

Many guardrail evaluations only check final response safety. But in real streaming generation, merely knowing "the final response has risk" is insufficient. A production-relevant question: when the model has generated a safe prefix but later slides into harmful advice, non-compliant steps, or unreliable conclusions, can the guardrail trigger at the right time? Trigger too late → risky content already shown to user; trigger too early → normal responses frequently interrupted.

This motivates SingStreamBench , released alongside. It doesn't just label whole responses "safe/unsafe"; samples are organized as "safe prefix + target continuation": guardrail should stay silent during safe prefix, trigger quickly after harmful content appears. SingStreamBench simultaneously examines three things:

Can it report: When risk truly appears, does the guardrail identify it?

Does it report too early: During safe prefix, does the guardrail stay restrained?

Is it timely: After risk appears, how long until the guardrail signals?

Core set contains 210 human-verified samples ; technical report also provides SingStreamBench-Full with 2,428 samples for broader evaluation. Dataset provides character-level onset of harmful content, enabling direct measurement of detection rate, false positive rate, premature trigger rate, and detection latency. This evaluation mirrors production: a qualified guardrail must both "report" and "report at the right time."

After Detecting Risk, How to Make the Model Correct Itself?

The value of endogenous risk signals goes beyond detection — they can become intervention signals during generation. The technical report extends this to medical generation with SingProbe-Med , decomposing generation-time safety control into: when to intervene, and how to intervene after triggering.

SingProbe-Med continuously perceives medical risk during generation, decides if/when to intervene; only when risk signal meets trigger condition does it activate risk-oriented sGDS decoding intervention , correcting local high-risk generation segments. Specifically, sGDS introduces an auxiliary branch trained on medical risk patterns as a negative reference; by contrasting main model vs risk-pattern branch generation tendencies, it suppresses candidate tokens more aligned with risk patterns, pulling local high-risk generation back toward safer directions. After intervention, model immediately resumes normal decoding.

SingProbe-Med intervention mechanism
SingProbe-Med intervention mechanism

Key: safety intervention doesn't accompany generation throughout, only intervenes when truly needed. In AntAngelMed-100B medical evaluation, full intervention corrected 25.03% of originally wrong baseline responses. Limiting single intervention to 64-token window retained 97.1% of full intervention effect.

Meanwhile, SingProbe-Med triggered intervention on only 4.69% of HealthBench samples; on AIME, CFBench, Arena-Hard-v2, AlignBench, WritingBench general capability benchmarks it never triggered. Most normal generation needs no extra correction flow.

On-demand intervention drastically cuts continuous dual-branch decoding cost. AntAngelMed-100B native single-branch decoding average request latency: 37.32s ; continuous dual-branch: ~ 60.93s ; with 64-token on-demand window: estimated 39.16s , only +1.84s (+4.94%) over native. SingProbe thus evolves from a continuous risk-score probe into a generation-time safety controller: idle normally, acts when risk appears.

Models and Deployment

From deployment perspective, SingProbe aims not to force users to rebuild safety services but to fit existing LLM inference pipelines. Each base model gets a corresponding SingProbe probe version. Probe head parameters: ~ 3–5M .

Released Models

inclusionAI/Ling-3.0-tiny-singprobe — base model: inclusionAI/Ling-3.0-tiny — applicable scenario: resource-constrained or low-latency services.

inclusionAI/Ling-3.0-flash-singprobe — base model: inclusionAI/Ling-3.0-flash — applicable scenario: higher detection capability service deployment.

Current release adapts Ling-3.0 series, GLM-5.2/5.3, Qwen, DeepSeekV4 — 29 mainstream open-source models — and integrates into SGLang and vLLM pipelines. Launch service specifying both base model and SingProbe checkpoint; generation output includes risk-score dictionary per output token.

python -m sglang.launch_server \
--model-path inclusionAI/Ling-3.0-flash \
--probe-ckpt inclusionAI/Ling-3.0-flash-singprobe \
--port 30000

This integration lets generated tokens and guardrail scores share the same inference pipeline, usable for real-time alerting, policy routing, generation termination, or subsequent constrained decoding. Support for more mainstream open-source models will be added continuously.

From External Add-on to Default Generation Signal

SingProbe transforms runtime guardrails from independent external audit modules into risk signals produced synchronously during model inference. Under this framework, user intent, response safety, and hallucination risk update continuously with token generation, directly serving downstream policies like alerting, refusal, termination, retry, and constrained decoding.

This release validates the endogenous design's feasibility: SingProbe matches or exceeds external baselines across multiple safety and reliability detection tasks while keeping decode overhead <0.5% ; SingStreamBench further evaluates risk trigger timing for streaming scenarios; SingProbe-Med demonstrates how risk signals enable generation-time intervention.

The goal is to provide a more generation-process-aligned implementation path for runtime guardrails, evolving safety capability from "optional external module" toward a default foundational capability in LLM deployment.

Open-Source Resources

Technical Report: https://arxiv.org/abs/2608.30703

Code & Integration: https://github.com/inclusionAI/SingProbe

Models: https://huggingface.co/collections/inclusionAI/singprobe

Benchmark: https://huggingface.co/datasets/inclusionAI/SingStreamBench

Developers, researchers, and industry partners are invited to try and provide feedback, jointly exploring more efficient and reliable generation-time safety mechanisms. For base model SingProbe adaptation requests, leave messages on GitHub/HuggingFace or email: [email protected], [email protected].

References

[1] OpenAI, "Model Development Cadence in the Era of Critical Cybersecurity Capabilities" , 2026-8-18.

[2] NVIDIA, "Measuring the Effectiveness and Performance of AI Guardrails in Generative AI Applications" , 2025.

[3] Zhang et al., "SoK: A Comprehensive Analysis and Evaluation of Guardrails for Large Language Models" , 2026.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

vLLMhallucination detectionSGLanghidden statesendogenous safetylightweight probesLLM guardrailsmedical AI safetyruntime guardrailsSingProbeSingProbe-MedSingStreamBenchstreaming detection
AntTech
Written by

AntTech

Technology is the core driver of Ant's future creation.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.