Tagged articles

Ascend NPU

9 articles · Page 1 of 1
Machine Heart
Machine Heart
Sep 11, 2026 · Artificial Intelligence

openJiuwen Launches Dual-Dimensional RSI Framework for Self-Improving AI Agents

openJiuwen introduces a dual-dimensional Recursive Self-Improvement (RSI) framework that enables AI agents to automatically optimize both their tooling (Harness) and deliverables (research papers, algorithms) on the WorkSwarm platform, with compute-affinity scheduling on Ascend NPUs cutting latency and resource usage, validated by SWE-bench pass-rate gains from 61% to 87%.

AI agentsAscend NPUCompute Affinity
0 likes · 15 min read
openJiuwen Launches Dual-Dimensional RSI Framework for Self-Improving AI Agents
Huawei Cloud Developer Alliance
Huawei Cloud Developer Alliance
Sep 8, 2026 · Artificial Intelligence

AReaL v1.0.5 LoRA RL: Low-Rank Adaptation for Accessible Large Model RL on Ascend

This article details AReaL-Ascend v1.0.5's LoRA RL capabilities, explaining how low-rank adaptation reduces memory overhead for large model reinforcement learning, describing two Megatron LoRA weight update modes (adapter sync vs. merge), and covering cross-node LoRA RL, MoE support, XCCL communication, and Qwen3.6-27B examples for practical deployment.

AReaLAscend NPULoRA
0 likes · 6 min read
AReaL v1.0.5 LoRA RL: Low-Rank Adaptation for Accessible Large Model RL on Ascend
Huawei Cloud Developer Alliance
Huawei Cloud Developer Alliance
Sep 6, 2026 · Artificial Intelligence

AReaL v1.0.5: Colocated Training/Inference & Multi-Teacher On-Policy Distillation for Efficient RL

AReaL-Ascend v1.0.5 introduces two major innovations: colocated training and inference on shared NPUs via Ray scheduling and AWEX IPC zero-copy, and Multi-Teacher On-Policy Distillation (MOPD) that fuses domain expert models into a single student using token-level teacher signals, demonstrated on Qwen3.6-27B with GRPO.

AReaLAWEXAscend NPU
0 likes · 12 min read
AReaL v1.0.5: Colocated Training/Inference & Multi-Teacher On-Policy Distillation for Efficient RL
Machine Heart
Machine Heart
Aug 11, 2026 · Artificial Intelligence

openJiuwen and Ascend Enable Agent “Compute‑Affinity”: Halve First‑Token Latency, Cut Inference Storage by 25%

The openJiuwen platform introduces a semantic coordination layer called Agent Hint, together with SAM and SPM managers, to align agent task states with compute resources, achieving a 57% reduction in first‑token latency, a 27.6% drop in end‑to‑end latency, a 33% increase in cache hit rate, and a 25% decrease in storage peak for multi‑agent inference workloads.

Agent HintAscend NPUCompute Affinity
0 likes · 11 min read
openJiuwen and Ascend Enable Agent “Compute‑Affinity”: Halve First‑Token Latency, Cut Inference Storage by 25%
Machine Heart
Machine Heart
Aug 8, 2026 · Artificial Intelligence

AURORA‑LM: Continuous‑Latent Diffusion Language Model on Ascend NPU

Researchers from Nanjing University’s PRLab present AURORA‑LM, a continuous‑latent diffusion language model trained on Ascend NPU up to 1 B parameters, detailing its two‑stage architecture, high‑capacity latent design, self‑trajectory consistency loss, and experimental gains over autoregressive and discrete diffusion baselines.

AURORA-LMAscend NPUcontinuous diffusion
0 likes · 10 min read
AURORA‑LM: Continuous‑Latent Diffusion Language Model on Ascend NPU
Huawei Cloud Developer Alliance
Huawei Cloud Developer Alliance
Apr 29, 2026 · Artificial Intelligence

Deploy DeepSeek‑V4 on Ascend NPU with Kthena in 3 Minutes (Prefill‑Decode Separation)

This guide walks through deploying the DeepSeek‑V4‑Flash model on Ascend NPU using Kthena’s ModelRoute, detailing the Prefill‑Decode (P/D) separation architecture, KV cache transfer via Mooncake, configuration of ModelServing and ModelRoute resources, and flexible scaling of Prefill and Decode replicas for optimal performance.

Ascend NPUDeepSeek-V4KV Cache
0 likes · 22 min read
Deploy DeepSeek‑V4 on Ascend NPU with Kthena in 3 Minutes (Prefill‑Decode Separation)
Huawei Cloud Developer Alliance
Huawei Cloud Developer Alliance
Apr 13, 2026 · Artificial Intelligence

How AReaL v1.0 Enables Scalable Agentic RL on Ascend NPU with AWEX Weight Sync

The new AReaL v1.0 release brings full Ascend NPU support, detailed installation guides, and a best‑practice example for training a 30B MoE model across four nodes, while the integrated AWEX weight‑sync mechanism dramatically reduces synchronization time, improving efficiency and stability for large‑scale Agentic RL workloads.

AWEXAgentic RLAscend NPU
0 likes · 12 min read
How AReaL v1.0 Enables Scalable Agentic RL on Ascend NPU with AWEX Weight Sync
AIWalker
AIWalker
Apr 13, 2025 · Artificial Intelligence

Huawei Pangu Ultra: 135B Ascend‑Native Dense LLM Without Nvidia GPUs

Huawei's Pangu Ultra introduces a 135‑billion‑parameter dense language model trained entirely on Ascend NPUs, detailing novel stability architectures, a domain‑aware tokenizer, multi‑stage pre‑training, extensive system optimizations, and benchmark results that surpass leading models such as Llama 405B and DeepSeek‑R1.

Ascend NPUDense ModelLarge Language Model
0 likes · 15 min read
Huawei Pangu Ultra: 135B Ascend‑Native Dense LLM Without Nvidia GPUs