Tagged articles

FP8 quantization

8 articles · Page 1 of 1
Machine Heart
Machine Heart
Sep 28, 2026 · Artificial Intelligence

SparkDiffusion: 265× Faster Video Generation on Consumer GPUs via Sparse Attention & Distillation

Researchers from Peking University, Tsinghua University, and Alibaba introduce SparkDiffusion, a unified framework combining sparse attention, few-step distillation, and FP8 quantization to achieve 265× acceleration for DiT video generation on a single RTX 5090, overcoming the high-sparsity trap that degrades quality at extreme sparsity.

DiTFP8 quantizationacceleration
0 likes · 12 min read
SparkDiffusion: 265× Faster Video Generation on Consumer GPUs via Sparse Attention & Distillation
Old Zhang's AI Learning
Old Zhang's AI Learning
Sep 10, 2026 · Artificial Intelligence

DeepSeek-V4.1-Flash: 8B Activated Model Beats 1.6T V4-Pro, Local Deployment Tested

DeepSeek-V4.1-Flash open-sourced with CED architecture and 8B/16B activated parameters outperforms its 1.6T predecessor V4-Pro on coding benchmarks, approaches GPT-6 Astra on DeepSWE, but requires 510GB FP8 weights needing 8×H200 for full-context local deployment; author tests across six harnesses finding Claude Code/Codex integration near peak performance.

AI model evaluationCED architectureDeepSWE
0 likes · 8 min read
DeepSeek-V4.1-Flash: 8B Activated Model Beats 1.6T V4-Pro, Local Deployment Tested
Old Zhang's AI Learning
Old Zhang's AI Learning
Jun 11, 2026 · Artificial Intelligence

Google’s 26B DiffusionGemma Model Delivers 1000+ Tokens/s – Runs on a 4090

DiffusionGemma, Google DeepMind’s 26B MoE model that generates 256‑token blocks via diffusion, achieves over 1000 tokens per second on H100/H200 GPUs, offers FP8 and NVFP4 quantized versions with near‑lossless accuracy, and can be deployed locally with vLLM Docker images, though it incurs higher first‑token latency and limited concurrency.

26B modelDiffusionGemmaFP8 quantization
0 likes · 10 min read
Google’s 26B DiffusionGemma Model Delivers 1000+ Tokens/s – Runs on a 4090
Old Zhang's AI Learning
Old Zhang's AI Learning
Apr 25, 2026 · Artificial Intelligence

Deploying DeepSeek‑V4‑Flash Locally on 2 × NVIDIA H20 (96 GB) – Quick Performance Test

This article walks through deploying DeepSeek‑V4‑Flash on a server with two NVIDIA H20 GPUs (96 GB each), detailing model download, Docker image preparation, launch script tweaks, memory compression via FP8 and expert parallelism, and reports observed concurrency limits and token‑per‑second speeds, including a test that disables the model's thinking mode.

DeepSeek-V4DockerFP8 quantization
0 likes · 6 min read
Deploying DeepSeek‑V4‑Flash Locally on 2 × NVIDIA H20 (96 GB) – Quick Performance Test
Old Zhang's AI Learning
Old Zhang's AI Learning
Apr 23, 2026 · Artificial Intelligence

DeepSeek Quietly Open‑Sources TileKernels to Push GPU Performance to Its Limits

DeepSeek has released TileKernels, a GPU kernel library written in the TileLang DSL, that targets H100/H200/B200 GPUs and claims to approach hardware limits in compute intensity and memory bandwidth, offering MoE routing, FP8/FP4 quantization, and dual‑language PyTorch references for deep‑learning engineers.

FP8 quantizationGPU optimizationLLM training
0 likes · 9 min read
DeepSeek Quietly Open‑Sources TileKernels to Push GPU Performance to Its Limits
Architects' Tech Alliance
Architects' Tech Alliance
May 2, 2025 · Artificial Intelligence

DeepSeek‑Prover‑V2‑671B: A Massive AI Model for Formal Mathematical Theorem Proving

DeepSeek‑Prover‑V2‑671B, a 671 billion‑parameter AI model released on Hugging Face, dramatically advances formal mathematical theorem proving with MoE architecture, FP8 quantization, 163 k token context, superior performance over GPT‑4 Turbo and other models, and broad implications for research and industry.

AIDeepSeekFP8 quantization
0 likes · 11 min read
DeepSeek‑Prover‑V2‑671B: A Massive AI Model for Formal Mathematical Theorem Proving