Tagged articles

Hardware‑software co‑design

7 articles · Page 1 of 1
Architects' Tech Alliance
Architects' Tech Alliance
Jul 20, 2026 · Artificial Intelligence

Supernode Architecture Explained: Definitions, Core Features, and Practical Use Cases

The whitepaper defines supernodes as high‑speed, tightly‑connected compute systems with unified memory addressing, microsecond‑level latency and terabyte‑per‑second bandwidth, outlines their physical, transaction, function and topology layers, demonstrates AI training and inference gains such as 80% communication reduction and 98.4% cluster scaling efficiency, and discusses industry impact, future scaling, standardization and green energy trends.

AI InfrastructureHardware‑software co‑designUnified memory
0 likes · 8 min read
Supernode Architecture Explained: Definitions, Core Features, and Practical Use Cases
Xiaomi Tech
Xiaomi Tech
May 18, 2026 · Artificial Intelligence

Xiaomi’s Imaging Algorithms Win CVPR 2026 NTIRE: Super‑Resolution, Portrait Restoration, Reflection Removal Breakthroughs

Xiaomi secured three top spots at CVPR 2026 NTIRE—first in Efficient Super‑Resolution with SPANV2, first in Portrait Restoration using a dual‑stage cascade, and second in Reflection Removal via RDNet‑XL and diffusion‑model distillation—showcasing hardware‑software co‑design, ultra‑fast inference, and novel algorithmic innovations.

Hardware‑software co‑designdiffusion model distillationefficient inference
0 likes · 14 min read
Xiaomi’s Imaging Algorithms Win CVPR 2026 NTIRE: Super‑Resolution, Portrait Restoration, Reflection Removal Breakthroughs
Machine Heart
Machine Heart
Apr 24, 2026 · Artificial Intelligence

Cambricon Achieves Day‑0 Native Support for DeepSeek‑V4, Uniting Two Chinese AI Leaders

Cambricon leveraged its NeuWare stack and vLLM framework to deliver Day‑0 native support for DeepSeek‑V4‑flash (285 B) and DeepSeek‑V4‑pro (1.6 T), open‑sourcing the adaptation and showcasing rapid model migration alongside extreme performance optimizations across software and hardware layers.

AI inferenceCambriconDeepSeek-V4
0 likes · 5 min read
Cambricon Achieves Day‑0 Native Support for DeepSeek‑V4, Uniting Two Chinese AI Leaders
AntTech
AntTech
Jul 18, 2025 · Artificial Intelligence

Explore the 2025 CCF‑Ant Research Fund: 50 Cutting‑Edge Projects in AI, Security & Computing

The CCF‑Ant Research Fund 2025, now open for its first batch, invites global university and institute researchers to apply by August 25 2025 for up to 50 projects spanning data security, hardware‑software co‑design, supercomputing, and artificial intelligence, with detailed topics, eligibility rules, and submission channels provided.

Hardware‑software co‑designResearch Fundingdata security
0 likes · 11 min read
Explore the 2025 CCF‑Ant Research Fund: 50 Cutting‑Edge Projects in AI, Security & Computing
Software Engineering 3.0 Era
Software Engineering 3.0 Era
May 15, 2025 · Artificial Intelligence

DeepSeek‑V3 Paper Reveals Breakthrough Hardware‑Software Co‑Design for AI Efficiency

DeepSeek‑V3 demonstrates that a tightly coupled hardware‑software design—featuring a memory‑saving MLA cache, a compute‑efficient DeepSeekMoE, a multi‑token prediction module, FP8 training, LogFMT compression, and an optimized eight‑plane fat‑tree network—can train a competitive LLM with only 2,048 H800 GPUs, cutting compute by up to 80% and boosting generation speed by 1.8×.

DeepSeek-V3FP8 trainingHardware‑software co‑design
0 likes · 12 min read
DeepSeek‑V3 Paper Reveals Breakthrough Hardware‑Software Co‑Design for AI Efficiency
Kuaishou Tech
Kuaishou Tech
Nov 8, 2021 · Artificial Intelligence

FPGA-Based Real-Time Streaming ASR Acceleration for Kuaishou: A Case Study in Domain-Specific Hardware Optimization

This paper presents a full fixed-point FPGA-based hardware acceleration solution for TDNN+LSTM acoustic models in real-time streaming ASR, achieving 37.67% latency reduction and 7.5x concurrency improvement through software-hardware co-design and domain-specific optimization.

Domain-specific ArchitectureFPGA accelerationHardware‑software co‑design
0 likes · 16 min read
FPGA-Based Real-Time Streaming ASR Acceleration for Kuaishou: A Case Study in Domain-Specific Hardware Optimization
Alibaba Cloud Developer
Alibaba Cloud Developer
May 21, 2019 · Artificial Intelligence

How Alibaba’s Offline AI Advances Model Compression and Edge Inference

Alibaba’s Machine Intelligence Lab shares two years of breakthroughs in offline AI, detailing low‑bit quantization, unified sparsity frameworks, hardware‑software co‑design, lightweight networks, and on‑device detection, alongside standardized training tools, multi‑platform inference engines, and productized edge solutions such as smart boxes and integrated cameras.

AIEdge InferenceHardware‑software co‑design
0 likes · 16 min read
How Alibaba’s Offline AI Advances Model Compression and Edge Inference