Tagged articles

diffusion models

187 articles · Page 1 of 2
Data Party THU
Data Party THU
Aug 18, 2026 · Artificial Intelligence

How to Master Online and Offline Policy Learning in Massive Action Spaces

This article reviews a PhD thesis that systematically studies online and offline learning for contextual bandits with huge action spaces, highlighting statistical, computational, and optimization challenges and presenting mixed‑effect Thompson sampling, diffusion priors, structured direct methods, and PAC‑Bayes pessimism as effective solutions.

PAC-BayesThompson samplingcontextual bandits
0 likes · 18 min read
How to Master Online and Offline Policy Learning in Massive Action Spaces
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Aug 13, 2026 · Artificial Intelligence

ARIS: Cross-Model Review and Persistent Memory Mechanisms for Reliable Long-Term Research Tasks

The talk introduces ARIS, an open‑source autonomous research system that uses cross‑model adversarial collaboration, a three‑layer evidence audit chain, and multi‑channel writing audit to ensure honest, end‑to‑end generation of research ideas through papers.

AI agentsARISautonomous research
0 likes · 4 min read
ARIS: Cross-Model Review and Persistent Memory Mechanisms for Reliable Long-Term Research Tasks
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Aug 12, 2026 · Artificial Intelligence

Achieving 4‑Step Diffusion Generation by Replacing MSE with Perceptual Loss in Five Lines of Code

By swapping the traditional MSE loss for a perceptual loss in Flow Matching training, the authors enable high‑quality diffusion generation in only 4–8 inference steps—down from 35–50—without teacher models, distribution or trajectory distillation, and they substantiate the claim with extensive experiments and a new distribution‑distance metric.

Distribution DistanceFew-Step GenerationFlow Matching
0 likes · 7 min read
Achieving 4‑Step Diffusion Generation by Replacing MSE with Perceptual Loss in Five Lines of Code
Machine Heart
Machine Heart
Aug 12, 2026 · Artificial Intelligence

Why Longer Captions Don't Boost Text‑to‑Image Models: Insights from ByteDance Seed

The ByteDance Seed team shows that merely extending caption length adds little visual supervision for diffusion models; instead, the amount of image‑grounded information in captions predicts training loss, leading them to propose Structured Prompt, new metrics (GPG, ED), and a three‑stage LLM prompter to improve both Diffusability and Promptability.

Effective DetailnessGPGcaption scaling
0 likes · 15 min read
Why Longer Captions Don't Boost Text‑to‑Image Models: Insights from ByteDance Seed
DeepHub IMBA
DeepHub IMBA
Aug 10, 2026 · Artificial Intelligence

Attention Heatmaps for Diffusion Models: Turning the Black‑Box into Explainability

This article explains how to compute and visualize attention heatmaps for text‑to‑image diffusion models, offering three complementary views (image‑to‑text, text‑to‑image, image‑to‑image), detailing the aggregation formulas, rendering process, and an interactive Flask web service that helps diagnose prompt failures and reveal model biases.

Flaskattention visualizationcross-attention
0 likes · 10 min read
Attention Heatmaps for Diffusion Models: Turning the Black‑Box into Explainability
Machine Heart
Machine Heart
Aug 9, 2026 · Artificial Intelligence

How to Train a One‑Step Generative Model Without CFG, DMD, GAN, or Drifting

The paper introduces TBSM, a lightweight direction‑tracking network that learns per‑sample Fake‑to‑Real vectors to guide a generator, enabling single‑forward (NFE=1) image synthesis with FID 1.92 on ImageNet‑512 and high‑quality text‑to‑image results, all without CFG, DMD, GAN, or drifting methods.

TBSMdiffusion modelsdirection tracking
0 likes · 11 min read
How to Train a One‑Step Generative Model Without CFG, DMD, GAN, or Drifting
Data Party THU
Data Party THU
Aug 9, 2026 · Artificial Intelligence

Breaking Scene Binding: Adaptive Diffusion Policy (DADP) Boosts Robot Generalization

Domain-Adaptive Diffusion Policy (DADP) decouples representation learning and injects domain information into the diffusion process, enabling robots to adapt across varying friction, mass, and dynamics, achieving strong zero-shot performance on MuJoCo and Adroit benchmarks, especially in out-of-distribution scenarios.

AdroitCross-Domain ControlDomain Adaptation
0 likes · 11 min read
Breaking Scene Binding: Adaptive Diffusion Policy (DADP) Boosts Robot Generalization
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Aug 7, 2026 · Artificial Intelligence

Real‑Time 16B‑Parameter Nano Banana Model Open‑Sourced for Video Editing

JD's JoyAI‑Video‑Edit brings a 16‑billion‑parameter, streaming‑capable AI model to real‑time video editing, achieving 30 FPS at 720p, beating prior streaming editors in speed, length handling, and benchmark scores while matching offline commercial quality.

AI video generationJoyAI-Video-Editbenchmark
0 likes · 15 min read
Real‑Time 16B‑Parameter Nano Banana Model Open‑Sourced for Video Editing
Bilibili Tech
Bilibili Tech
Jul 24, 2026 · Artificial Intelligence

Geometry‑Aware, Training‑Free Acceleration of Diffusion Transformer Sampling (CVPR 2026 Highlight)

GeoRK2 introduces a training‑free, plug‑and‑play framework that combines second‑order Runge‑Kutta integration with low‑rank geometric correction of diffusion Transformers, enabling 4‑5× faster image, text‑to‑image, and video generation while preserving high visual quality and stability.

Geometry-Aware SamplingPlug-and-PlayRunge-Kutta
0 likes · 17 min read
Geometry‑Aware, Training‑Free Acceleration of Diffusion Transformer Sampling (CVPR 2026 Highlight)
DaTaobao Tech
DaTaobao Tech
Jul 10, 2026 · Artificial Intelligence

Rethinking Diffusion‑Based Video Super‑Resolution with Dense Feature‑Guided Alignment (DGAF‑VSR)

The paper introduces DGAF‑VSR, a diffusion‑model video super‑resolution framework that leverages feature‑domain alignment and dense temporal guidance via an Optical‑Guided Warping Module and a Feature‑wise Temporal Condition Module, achieving state‑of‑the‑art perceptual, fidelity, and temporal scores on REDS4, Vid4 and VideoLQ datasets.

CVPR 2026DGAF-VSRdiffusion models
0 likes · 12 min read
Rethinking Diffusion‑Based Video Super‑Resolution with Dense Feature‑Guided Alignment (DGAF‑VSR)
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 6, 2026 · Artificial Intelligence

ICML 2026 Opens – Tsinghua Wins Outstanding Paper, DeepMind Earns Test‑of‑Time Award, and Who Is Machine Learning For?

ICML 2026 in Seoul broke submission records, sparked controversy over LLM‑generated reviews, honored breakthrough papers on diffusion models, reinforcement learning and AI alignment, and culminated in a reflective question about the true purpose and beneficiaries of machine learning.

AI ethicsGrokkingICML 2026
0 likes · 15 min read
ICML 2026 Opens – Tsinghua Wins Outstanding Paper, DeepMind Earns Test‑of‑Time Award, and Who Is Machine Learning For?
Machine Heart
Machine Heart
Jul 6, 2026 · Artificial Intelligence

ICML 2026 Awards Unveiled: Breakthroughs in Diffusion Models, AI Alignment, and Reinforcement Learning

ICML 2026 announced ten award‑winning papers, highlighting novel insights such as the flexibility trap in diffusion language models, high‑accuracy sampling for diffusion, risks of AI alignment tools, a random‑matrix view of diffusion consistency, grokking in ridge regression, and an asynchronous deep‑RL framework, each accompanied by concise abstracts and links.

AI AlignmentGrokkingICML 2026
0 likes · 16 min read
ICML 2026 Awards Unveiled: Breakthroughs in Diffusion Models, AI Alignment, and Reinforcement Learning
Machine Heart
Machine Heart
Jul 2, 2026 · Artificial Intelligence

EMCES: How Episodic Memory Guides Controllable Sample Synthesis to Boost Reinforcement Learning

The paper introduces EMCES, a method that injects episodic memory into controllable diffusion models and uses a hash‑based state representation to generate high‑value synthetic samples, dramatically improving sample efficiency and downstream reinforcement‑learning performance while cutting storage and time costs.

Episodic MemoryHashingOffline RL
0 likes · 14 min read
EMCES: How Episodic Memory Guides Controllable Sample Synthesis to Boost Reinforcement Learning
vivo Internet Technology
vivo Internet Technology
Jul 1, 2026 · Artificial Intelligence

LearnIR: Posterior Sampling for Image Restoration – Face Shadow Removal & Dehazing (ICLR 2026)

LearnIR tackles real‑world image restoration under heterogeneous degradations by training a lightweight network to predict gradient‑correction distributions for diffusion posterior sampling without a forward operator, and adds a dynamic‑resolution module that further suppresses noise, achieving state‑of‑the‑art PSNR, SSIM and LPIPS on multiple benchmarks.

LearnIRdehazingdiffusion models
0 likes · 10 min read
LearnIR: Posterior Sampling for Image Restoration – Face Shadow Removal & Dehazing (ICLR 2026)
Machine Heart
Machine Heart
Jun 29, 2026 · Artificial Intelligence

Control Humanoid Robot Motion with a Sentence or Music via OMG Framework

OMG introduces a hierarchical “generation brain + tracking cerebellum” framework that leverages a large multimodal dataset and diffusion‑based OMG‑DiT network to let humanoid robots synthesize full‑body motions from a single sentence, music clip, or pose, achieving state‑of‑the‑art performance across text, audio, and motion benchmarks.

AI generationHumanoid RoboticsOMG framework
0 likes · 11 min read
Control Humanoid Robot Motion with a Sentence or Music via OMG Framework
Data Party THU
Data Party THU
Jun 23, 2026 · Artificial Intelligence

How Diffusion Models Achieve Generalization: Insights from a CVPR 2026 Tutorial

Diffusion models have set the state‑of‑the‑art in image, video, and audio generation, yet their training objective admits a unique closed‑form solution that merely memorizes training data; this tutorial examines why they still generalize by exploring score smoothing, architectural inductive bias, training dynamics, and data geometry, all illustrated with hands‑on Jupyter notebooks.

CVPR 2026Generative Modelingdata geometry
0 likes · 2 min read
How Diffusion Models Achieve Generalization: Insights from a CVPR 2026 Tutorial
AI Architecture Hub
AI Architecture Hub
Jun 23, 2026 · Artificial Intelligence

Top AI Papers This Week (June 14‑21): SpatialClaw, SkillWeaver, PreAct, and More

This article reviews seven recent AI research papers, detailing how SpatialClaw enables code‑based spatial reasoning for vision‑language models, SkillWeaver introduces compositional skill routing, PreAct compiles agent actions into reusable state‑machines, and other works advance world‑model inference, self‑designing RL environments, collective skill‑tree search, and process‑aligned reinforcement learning for diffusion LLMs.

Large Language Modelsagent reasoningdiffusion models
0 likes · 15 min read
Top AI Papers This Week (June 14‑21): SpatialClaw, SkillWeaver, PreAct, and More
Amap Tech
Amap Tech
Jun 22, 2026 · Artificial Intelligence

Three Amap Papers Accepted at IROS 2026: VLN Navigation, VLA Latency Correction, and Diffusion‑Based Quadruped Control

IROS 2026 received 4,348 submissions and accepted 1,585 papers (36% acceptance); Amap had three papers selected, covering online semantic‑affordance navigation, an asynchronous edge adapter for VLA‑based navigation, and a diffusion‑guided constraint‑aware framework for high‑fidelity quadruped locomotion.

asynchronous VLA navigationdiffusion modelsembodied AI
0 likes · 8 min read
Three Amap Papers Accepted at IROS 2026: VLN Navigation, VLA Latency Correction, and Diffusion‑Based Quadruped Control
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jun 18, 2026 · Artificial Intelligence

UniRL: Tencent Hunyuan’s Open‑Source Framework Unifying Multimodal RL Training

UniRL is an open‑source, distributed reinforcement‑learning post‑training framework that consolidates fragmented pipelines for image, video, and language‑vision models, offering a unified rollout‑reward‑advantage‑train‑sync contract, extensive model support, built‑in algorithms, and multi‑modal reward components to lower engineering barriers in AIGC research.

Distributed TrainingLLMMultimodal RL
0 likes · 10 min read
UniRL: Tencent Hunyuan’s Open‑Source Framework Unifying Multimodal RL Training
Machine Heart
Machine Heart
Jun 10, 2026 · Artificial Intelligence

DRDD: Turning Diffusion Noise into a Domain Harmonizer for Image Translation

The paper introduces Decoupled Residual Denoising Diffusion (DRDD), which reinterprets Gaussian noise as a domain harmonizer and separates residual removal from denoising, enabling more data‑efficient, multi‑task image‑to‑image translation and achieving state‑of‑the‑art results on benchmarks such as All‑in‑One‑5 with limited paired data.

DRDDData Efficiencycomputer vision
0 likes · 14 min read
DRDD: Turning Diffusion Noise into a Domain Harmonizer for Image Translation
Machine Heart
Machine Heart
May 29, 2026 · Artificial Intelligence

WaDi: One‑Step Image Generation with LoRA Meets RoPE

This work analyzes weight‑direction changes in diffusion‑model distillation, proposes a low‑rank rotation adapter (LoRaD) to model those changes, and integrates it into Variational Score Distillation as WaDi, achieving state‑of‑the‑art FID on COCO with only ~10% trainable parameters while generalizing to multiple downstream tasks.

LoRARoPEdiffusion models
0 likes · 20 min read
WaDi: One‑Step Image Generation with LoRA Meets RoPE
Machine Heart
Machine Heart
May 29, 2026 · Artificial Intelligence

DiffusionOPD: A New Online Policy Distillation Paradigm for Multi‑Task Diffusion Models

DiffusionOPD introduces a unified on‑policy distillation framework for diffusion models that decouples single‑task online policy exploration from multi‑task capability integration, training expert teachers per task and distilling their skills into a single student model, achieving faster convergence and higher performance across composition, OCR, and aesthetic tasks.

KL DivergenceOn-Policy DistillationPPO
0 likes · 8 min read
DiffusionOPD: A New Online Policy Distillation Paradigm for Multi‑Task Diffusion Models
Machine Heart
Machine Heart
May 25, 2026 · Artificial Intelligence

VeRL-Omni: Universal RL Post‑Training for Diffusion and Multimodal Models

VeRL-Omni is an open‑source RL post‑training framework built on verl and vLLM‑Omni that enables efficient, high‑throughput rollout and flexible reward computation for diffusion, AR‑DiT, and unified multimodal generation models, supporting diverse hardware, modular trainers, and demonstrating up to 14% latency reduction and high training throughput in benchmark experiments.

FlowGRPORLVeRL-Omni
0 likes · 9 min read
VeRL-Omni: Universal RL Post‑Training for Diffusion and Multimodal Models
Machine Heart
Machine Heart
May 25, 2026 · Artificial Intelligence

Breaking the Reward Trade‑off: Flow‑OPD Brings Multi‑Teacher OPD to Image Generation

Flow‑OPD introduces on‑policy distillation into flow‑matching diffusion models, using a multi‑teacher online rollout framework and manifold‑anchor regularization to resolve the seesaw effect of single and mixed rewards, achieving superior multi‑task performance and surpassing specialist models in image generation.

Flow-OPDManifold Anchor RegularizationOn-Policy Distillation
0 likes · 9 min read
Breaking the Reward Trade‑off: Flow‑OPD Brings Multi‑Teacher OPD to Image Generation
Machine Heart
Machine Heart
May 24, 2026 · Artificial Intelligence

How Hallo‑Live Achieves Real‑Time Streaming Text‑Driven Audio‑Video Avatar Generation

Hallo‑Live introduces an asynchronous dual‑stream diffusion framework combined with human‑centric preference‑guided distillation, enabling text‑driven audio‑video avatars to run at 20.38 FPS with 0.94 s latency—over 16× faster and 99.3 % lower latency than the teacher Ovi model while preserving visual quality and lip‑sync.

Hallo-LiveNVIDIA H200asynchronous dual-stream diffusion
0 likes · 9 min read
How Hallo‑Live Achieves Real‑Time Streaming Text‑Driven Audio‑Video Avatar Generation
Machine Heart
Machine Heart
May 21, 2026 · Artificial Intelligence

RAEv2: How a Simple Extra Operation Makes Image Generation Train Ten Times Faster

The RAEv2 framework replaces traditional VAEs by summing multiple layers of pretrained vision encoders, combines RAE with REPA for complementary semantic and spatial gains, and leverages free guidance, achieving up to ten‑fold faster convergence, higher image quality, and lower compute on ImageNet‑256 diffusion training.

RAEv2Representation AutoencoderVision Encoders
0 likes · 11 min read
RAEv2: How a Simple Extra Operation Makes Image Generation Train Ten Times Faster
Machine Heart
Machine Heart
May 11, 2026 · Artificial Intelligence

UniVidX Sets New SOTA on Multiple Video Tasks – A Unified Multimodal Framework Presented at SIGGRAPH 2026

UniVidX, a unified multimodal framework for video generation and understanding accepted at SIGGRAPH 2026, reformulates diverse video graphics tasks as conditional generation, achieving or surpassing state‑of‑the‑art performance while demonstrating strong data efficiency and cross‑domain generalization.

Data EfficiencyMultimodal Video GenerationSIGGRAPH 2026
0 likes · 10 min read
UniVidX Sets New SOTA on Multiple Video Tasks – A Unified Multimodal Framework Presented at SIGGRAPH 2026
Machine Heart
Machine Heart
May 8, 2026 · Artificial Intelligence

Omni2Sound Beats Multi-Modal Audio ‘Generalist’ Dilemma via Data Alignment

Omni2Sound tackles the long‑standing “generalist” dilemma of unified audio generation by constructing a high‑quality V‑T‑A dataset (SoundAtlas), employing a three‑stage progressive training pipeline, and using a simple Diffusion Transformer backbone, ultimately achieving state‑of‑the‑art performance on T2A, V2A and VT2A tasks and strong robustness on off‑screen scenarios.

Data AlignmentOmni2SoundTask Competition
0 likes · 16 min read
Omni2Sound Beats Multi-Modal Audio ‘Generalist’ Dilemma via Data Alignment
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Apr 27, 2026 · Artificial Intelligence

From Parameter Tuning to Control: CFG‑Ctrl Boosts Stability and Precision in Text‑to‑Image Generation

The paper introduces CFG‑Ctrl, a control‑theoretic redesign of classifier‑free diffusion guidance that treats the generation process as a dynamic system, achieving more stable and accurate text‑to‑image results across multiple model scales and evaluation metrics.

CFG-CtrlControl Theoryclassifier-free guidance
0 likes · 15 min read
From Parameter Tuning to Control: CFG‑Ctrl Boosts Stability and Precision in Text‑to‑Image Generation
Kuaishou Tech
Kuaishou Tech
Apr 24, 2026 · Artificial Intelligence

ICLR 2026: Kuaishou Tech Team’s Cutting‑Edge AI Research Highlights

This article reviews eight Kuaishou‑authored papers accepted at ICLR 2026, summarizing their problem statements, novel methods such as front‑door causal attribution, visual table retrieval, denoising rerankers, difficulty‑adaptive reasoning, diffusion code infilling, generative ordinal regression, multimodal video retrieval, e‑commerce dialogue benchmarks, and a new LLM creativity evaluator, together with reported experimental gains.

Artificial IntelligenceICLR 2026Kuaishou
0 likes · 19 min read
ICLR 2026: Kuaishou Tech Team’s Cutting‑Edge AI Research Highlights
SuanNi
SuanNi
Apr 21, 2026 · Artificial Intelligence

Why AI Video Generation Is Leaving the Silent Era: Architecture, Alignment, and Evaluation Insights

This article analyzes the rapid evolution of multimodal video generation models from separated visual‑audio pipelines to unified diffusion Transformers, detailing VAE compression, MoE scaling, cross‑modal alignment techniques, comprehensive evaluation metrics, real‑world applications, and the remaining technical challenges.

Video Generationaudio-visual alignmentdiffusion models
0 likes · 15 min read
Why AI Video Generation Is Leaving the Silent Era: Architecture, Alignment, and Evaluation Insights
AI Explorer
AI Explorer
Apr 16, 2026 · Artificial Intelligence

AI Tech Daily: Top AI Research and Industry Updates on April 16 2026

This roundup highlights recent AI breakthroughs such as NVIDIA‑MIT’s Sol‑RL framework for faster diffusion model training, Peking University’s CPL++ visual localization improvement, DeepMind’s TIPSv2 for image recognition, Boston Dynamics Spot’s AI upgrade, Anthropic’s safety paper, a major MCP protocol vulnerability, OpenAI’s GPT‑5.4 release, and the shifting AI video landscape.

AIAI safetyLarge Language Models
0 likes · 5 min read
AI Tech Daily: Top AI Research and Industry Updates on April 16 2026
AI Explorer
AI Explorer
Apr 16, 2026 · Artificial Intelligence

How NVIDIA, HKU, and MIT’s Sol‑RL Framework Supercharges Diffusion Model Training

NVIDIA, Hong Kong University, and MIT introduced the Sol‑RL framework, which uses reinforcement‑learning‑guided sampling to cut diffusion model training time by several‑fold without sacrificing image quality, potentially lowering entry barriers for small teams and shifting the AIGC industry toward an efficiency‑driven competition.

AIGCNvidiaSol-RL
0 likes · 6 min read
How NVIDIA, HKU, and MIT’s Sol‑RL Framework Supercharges Diffusion Model Training
Machine Heart
Machine Heart
Apr 16, 2026 · Artificial Intelligence

Achieving 4.6× Faster Diffusion Model Training with FP4‑BF16 Dual‑Track Parallelism (Sol‑RL)

Sol‑RL, a framework from NVIDIA, Hong Kong University and MIT, integrates NVFP4 inference for large‑scale rollout exploration and BF16 precision for high‑fidelity regeneration, delivering up to 4.64× faster convergence at equivalent reward levels while preserving BF16 training fidelity across SANA, FLUX.1 and SD3.5‑L models.

BF16FP4GPU Optimization
0 likes · 9 min read
Achieving 4.6× Faster Diffusion Model Training with FP4‑BF16 Dual‑Track Parallelism (Sol‑RL)
SuanNi
SuanNi
Apr 12, 2026 · Artificial Intelligence

How TDM‑R1 Achieves 4‑Step Image Generation that Beats 80‑Step Models

Researchers from HKUST, CUHK and XiaoHongShu introduced TDM‑R1, a reinforcement‑learning‑based method that enables 4‑step diffusion image generation to surpass 80‑step models in speed, fidelity, and complex instruction adherence, as demonstrated on the GenEval benchmark and multiple quality metrics.

AI image synthesisBenchmarkingFew-Step Generation
0 likes · 9 min read
How TDM‑R1 Achieves 4‑Step Image Generation that Beats 80‑Step Models
Bighead's Algorithm Notes
Bighead's Algorithm Notes
Apr 9, 2026 · Artificial Intelligence

WSDM2026 Quantitative Research Papers: Summaries and Insights

This article presents concise summaries of three recent AI‑driven finance papers—Diffolio’s diffusion‑based risk‑aware portfolio optimization, STORM’s dual‑vector‑quantized VAE factor model, and AutoHypo‑Fin’s autonomous web‑mined hypothesis generation—highlighting their motivations, methods, and experimental gains.

AI for financeVQ-VAEdiffusion models
0 likes · 9 min read
WSDM2026 Quantitative Research Papers: Summaries and Insights
HyperAI Super Neural
HyperAI Super Neural
Apr 7, 2026 · Artificial Intelligence

MIT’s DRiffusion Achieves 1.4–3.7× Faster Diffusion Sampling via Draft‑and‑Refine Parallelism

MIT researchers introduce DRiffusion, a draft‑and‑refine parallel framework that uncovers intrinsic parallelism in diffusion models, delivering 1.4–3.7× speedup on three GPUs while preserving near‑lossless image quality across Stable Diffusion 2.1, SDXL and SD3 evaluated on MS‑COCO.

AI accelerationDRiffusionMS-COCO
0 likes · 14 min read
MIT’s DRiffusion Achieves 1.4–3.7× Faster Diffusion Sampling via Draft‑and‑Refine Parallelism
vivo Internet Technology
vivo Internet Technology
Apr 1, 2026 · Artificial Intelligence

Why Fixed CFG Fails and How Time‑Adaptive C²FG Boosts Diffusion Image Generation

This article introduces C²FG, a training‑free, plug‑and‑play time‑adaptive exponential control function that replaces the fixed classifier‑free guidance scale, theoretically justifies its superiority with score discrepancy bounds, and demonstrates significant FID and IS improvements across multiple diffusion architectures on ImageNet.

CVPR 2026Plug-and-Playclassifier-free guidance
0 likes · 7 min read
Why Fixed CFG Fails and How Time‑Adaptive C²FG Boosts Diffusion Image Generation
Bighead's Algorithm Notes
Bighead's Algorithm Notes
Mar 31, 2026 · Artificial Intelligence

Top AI-Driven Quantitative Finance Papers from AAAI 2026

This article curates and summarizes recent AI research papers presented at AAAI 2026 that advance quantitative finance, covering controllable market generation, LLM‑powered alpha factor mining, risk‑aware multi‑agent portfolio management, foundation models for market data, and reinforcement‑learning trading policies.

AIFinancial Market Simulationdiffusion models
0 likes · 12 min read
Top AI-Driven Quantitative Finance Papers from AAAI 2026
PaperAgent
PaperAgent
Mar 28, 2026 · Artificial Intelligence

How ACCORD Breaks Concept Coupling in Custom Text‑to‑Image Generation

The ACCORD framework formalizes the concept‑coupling issue in text‑to‑image diffusion models as a statistical dependency problem and resolves it with two plug‑and‑play regularization losses, dramatically improving fidelity and text control without altering model architecture.

ACCORDAI researchconcept coupling
0 likes · 7 min read
How ACCORD Breaks Concept Coupling in Custom Text‑to‑Image Generation
HyperAI Super Neural
HyperAI Super Neural
Mar 25, 2026 · Artificial Intelligence

Low‑Barrier Deployment of NVIDIA’s Latest Physical AI Models for Humanoid Robots, Motion Generation, and Diffusion Fine‑Tuning

The article introduces NVIDIA’s Physical AI suite announced at GTC 2026—including Isaac GR00T, SOMA‑X, Kimodo, and FDFO—explains each model’s architecture and purpose, and provides one‑click online tutorials that let developers experiment with humanoid robotics, human‑body modeling, motion generation, and diffusion model fine‑tuning at minimal cost.

FDFOIsaac GR00TKimodo
0 likes · 8 min read
Low‑Barrier Deployment of NVIDIA’s Latest Physical AI Models for Humanoid Robots, Motion Generation, and Diffusion Fine‑Tuning
Bighead's Algorithm Notes
Bighead's Algorithm Notes
Mar 22, 2026 · Artificial Intelligence

DigMA: Controllable Generation of Financial Market Data – A Deep Dive

This article reviews the DigMA model, which uses a diffusion‑guided meta‑agent to generate high‑fidelity, controllable order‑flow data for financial markets, details its problem formulation, architecture, training on Chinese stock datasets, extensive experiments—including reinforcement‑learning‑based high‑frequency trading evaluation—and demonstrates its superior accuracy and ultra‑low latency generation.

Financial Market SimulationMeta‑Agentcontrollable generation
0 likes · 16 min read
DigMA: Controllable Generation of Financial Market Data – A Deep Dive
Bighead's Algorithm Notes
Bighead's Algorithm Notes
Mar 13, 2026 · Artificial Intelligence

Paper Reading: STABLE – A Robust Portfolio Allocation Method Using Conditional Diffusion Estimates

The STABLE framework integrates a conditional diffusion generator with a Black‑Litterman mean‑variance optimizer to produce style‑aware return forecasts and risk‑aware portfolio weights, achieving up to a 122.9% Sharpe‑ratio boost, lower drawdowns, and a 15.7% MSE reduction across major equity markets.

Black-LittermanFinancial AIconditional diffusion
0 likes · 17 min read
Paper Reading: STABLE – A Robust Portfolio Allocation Method Using Conditional Diffusion Estimates
AIWalker
AIWalker
Mar 10, 2026 · Artificial Intelligence

MIGM-Shortcut: Learning Controlled Latent Dynamics to Speed Up Masked Image Generation

The paper introduces MIGM-Shortcut, a self‑supervised method that learns controlled latent‑state dynamics to bypass redundant bidirectional attention in Masked Image Generation Models, achieving over 4× speed‑up on state‑of‑the‑art multimodal diffusion models like Lumina‑DiMOO while preserving image quality.

AIMIGMdiffusion models
0 likes · 8 min read
MIGM-Shortcut: Learning Controlled Latent Dynamics to Speed Up Masked Image Generation
SuanNi
SuanNi
Feb 23, 2026 · Artificial Intelligence

How FireRed-Image-Edit Sets New Standards for AI-Powered Image Editing

FireRed-Image-Edit, an open‑source instruction‑driven diffusion model, combines massive high‑quality data, a dual‑stream multimodal architecture, progressive training, and a comprehensive multi‑dimensional benchmark to achieve unprecedented pixel‑level control and human‑like editing performance across diverse visual tasks.

AIData EngineeringTraining Strategies
0 likes · 12 min read
How FireRed-Image-Edit Sets New Standards for AI-Powered Image Editing
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Feb 14, 2026 · Artificial Intelligence

Latent Forcing: Reordering Diffusion Steps Boosts Pixel‑Level Image Quality

The new Latent Forcing technique from Fei‑Fei Li’s team reorders the diffusion trajectory, first generating a latent structural sketch and then refining pixel details, which restores efficiency of latent‑space models while preserving 100 % pixel fidelity, achieving state‑of‑the‑art FID scores on ImageNet‑256.

AI researchImageNetdiffusion models
0 likes · 6 min read
Latent Forcing: Reordering Diffusion Steps Boosts Pixel‑Level Image Quality
Design Hub
Design Hub
Jan 17, 2026 · Artificial Intelligence

FLUX.2 Klein Generates Images in Under a Second and Unlocks Midjourney‑Style Prompts

The article reviews Black Forest Labs' FLUX.2 Klein model, highlighting its sub‑second 1024×1024 image generation, low‑VRAM requirements, four‑step inference speedups, and competitive quality versus SD3 and Midjourney V6, while also sharing Midjourney‑style prompt examples for creative design.

AI image generationFLUX.2GPU Acceleration
0 likes · 8 min read
FLUX.2 Klein Generates Images in Under a Second and Unlocks Midjourney‑Style Prompts
Design Hub
Design Hub
Dec 22, 2025 · Artificial Intelligence

Open‑Source AI Photoshop: Alibaba’s Qwen‑Image‑Layered Enables One‑Click Smart Layering

Alibaba’s Qwen‑Image‑Layered model, now fully open‑source, automatically separates a single image into editable RGBA layers using diffusion, offering Photoshop‑level editing, prompt‑controlled layer counts, and deep decomposition, with applications ranging from PPT de‑construction to game asset extraction, while noting limitations on realistic photos.

AI image segmentationComfyUIFigma plugin
0 likes · 8 min read
Open‑Source AI Photoshop: Alibaba’s Qwen‑Image‑Layered Enables One‑Click Smart Layering
Data Party THU
Data Party THU
Dec 18, 2025 · Artificial Intelligence

How Diffusion Models and Transformers Power the Next Generation of AI Video Generation

AI video generation now turns textual prompts into high‑quality clips using diffusion models and transformer‑based architectures; this article explains the underlying mathematics, training objectives, spatio‑temporal encoding, breakthroughs like consistent motion and physical realism, and discusses the technology’s opportunities and inherent risks.

AI video generationSpatio-temporal modelingTransformers
0 likes · 11 min read
How Diffusion Models and Transformers Power the Next Generation of AI Video Generation
Alibaba Cloud Developer
Alibaba Cloud Developer
Dec 18, 2025 · Artificial Intelligence

How to Build a Real‑Time AI‑Powered Anime‑Style Video Generator for Social Apps

This technical report details the end‑to‑end workflow for integrating an AIGC video generation module into a social app, covering requirement analysis, model and hardware selection, dataset construction, LoRA and full‑parameter training, multiple acceleration techniques such as Sage Attention, TeaCache, XDiT, gradient‑checkpointing offload, tiled VAE, and quantization, followed by extensive performance evaluation and metric‑based ranking of the final models.

AI video generationLoRA fine-tuningModel Optimization
0 likes · 38 min read
How to Build a Real‑Time AI‑Powered Anime‑Style Video Generator for Social Apps
Kuaishou Tech
Kuaishou Tech
Dec 3, 2025 · Artificial Intelligence

Can Diffusion Models Be Their Own Reward Model? Latent Reward Modeling & Step-Level Preference Optimization

This article presents a novel paradigm—Latent Reward Model (LRM) and Latent Preference Optimization (LPO)—that repurposes diffusion models as noise‑aware latent reward models for step‑level preference optimization, addressing the shortcomings of pixel‑level reward models, introducing multi‑preference consistent filtering, and demonstrating significant performance and efficiency gains on benchmarks such as PickScore and T2I‑CompBench++.

AI Alignmentdiffusion modelsimage generation
0 likes · 9 min read
Can Diffusion Models Be Their Own Reward Model? Latent Reward Modeling & Step-Level Preference Optimization
HyperAI Super Neural
HyperAI Super Neural
Nov 19, 2025 · Artificial Intelligence

LocDiff: Achieving Global-Scale Precise Image Geolocation Without Grids or Reference Libraries

The LocDiff framework introduces a spherical‑harmonics Dirac‑delta encoding and a conditional Siren‑UNet diffusion model that enables accurate worldwide image geolocation without relying on predefined grids or external image libraries, outperforming prior methods in precision, generalization, and computational efficiency.

AI researchLocDiffdiffusion models
0 likes · 16 min read
LocDiff: Achieving Global-Scale Precise Image Geolocation Without Grids or Reference Libraries
Bighead's Algorithm Notes
Bighead's Algorithm Notes
Nov 15, 2025 · Artificial Intelligence

Quantitative Finance Paper Digest: Nov 8‑14 2025 Highlights

This article summarizes five recent arXiv papers that apply advanced AI techniques such as diffusion models, hierarchical attention, and stochastic differential equations to multivariate financial time‑series forecasting, portfolio selection, volatility surface generation, and gold‑futures alpha strategies, presenting their core methods and experimental results.

diffusion modelsequilibrium portfoliofinancial time series
0 likes · 10 min read
Quantitative Finance Paper Digest: Nov 8‑14 2025 Highlights
Kuaishou Tech
Kuaishou Tech
Nov 13, 2025 · Artificial Intelligence

Unlocking Unusual Concept Combinations in Generative AI with IMBA Loss

The paper identifies imbalanced concept distributions as the main obstacle to arbitrary concept‑combination in text‑to‑image/video generation, proposes the token‑level IMBA Distance and a lightweight IMBA Loss that adaptively re‑weights training tokens, and demonstrates through extensive experiments and a new Inert‑CompBench benchmark that this loss dramatically improves compositional ability without extra data.

IMBA Lossbenchmarkconcept combination
0 likes · 9 min read
Unlocking Unusual Concept Combinations in Generative AI with IMBA Loss
AI Frontier Lectures
AI Frontier Lectures
Nov 4, 2025 · Artificial Intelligence

How DiffPathV2 Achieves Zero‑Shot Image Anomaly Detection with 94.9% AUROC

This article breaks down the ICCV 2025 paper "Zero‑Shot Image Anomaly Detection Using Generative Foundation Models," explaining how DiffPathV2 leverages diffusion model denoising trajectories, six‑dimensional score errors, and SSIM weighting to detect out‑of‑distribution images without any task‑specific training, achieving state‑of‑the‑art AUROC scores across multiple benchmarks.

AUROCDiffPathV2SSIM
0 likes · 10 min read
How DiffPathV2 Achieves Zero‑Shot Image Anomaly Detection with 94.9% AUROC
Data Party THU
Data Party THU
Oct 15, 2025 · Artificial Intelligence

Designing Safe, Sample-Efficient, and Robust Reinforcement Learning for Ranking and Diffusion Models

This paper proposes a reinforcement‑learning framework that simultaneously ensures safety, sample efficiency, and robustness, applying a contextual‑bandit perspective to ranking/recommendation systems and text‑to‑image diffusion models, and introduces novel algorithms for safe deployment, variance‑reduced off‑policy estimation, and a LOOP method for generative RL.

RobustnessSafetycontextual bandits
0 likes · 5 min read
Designing Safe, Sample-Efficient, and Robust Reinforcement Learning for Ranking and Diffusion Models
AI Algorithm Path
AI Algorithm Path
Oct 15, 2025 · Artificial Intelligence

Building a Flow Matching Model from Scratch: Theory Explained

This article walks through the theory behind flow‑matching generative models, contrasting them with diffusion models, detailing the velocity‑field formulation, training objective, and sampling procedure, and includes visual illustrations of the core concepts.

Flow MatchingGenerative ModelsODE
0 likes · 8 min read
Building a Flow Matching Model from Scratch: Theory Explained
Data Party THU
Data Party THU
Oct 13, 2025 · Artificial Intelligence

How BranchGRPO Accelerates and Stabilizes Diffusion Model Alignment

BranchGRPO introduces a tree‑structured branching, reward‑fusion, and lightweight pruning framework that dramatically speeds up diffusion and flow model training while delivering denser, more stable reward signals, achieving up to five‑fold faster convergence and higher alignment scores on image and video generation benchmarks.

BranchGRPOEfficiencyRLHF
0 likes · 10 min read
How BranchGRPO Accelerates and Stabilizes Diffusion Model Alignment
AI Algorithm Path
AI Algorithm Path
Oct 12, 2025 · Artificial Intelligence

Flow Matching vs Diffusion Models: Key Differences and Connections

This technical article provides a comprehensive comparison of diffusion models and flow matching, covering their intuitive explanations, underlying mathematics, training objectives, sampling efficiency, theoretical guarantees, practical examples, and code implementations to illustrate how each generative approach works.

Flow Matchingdiffusion modelsgenerative AI
0 likes · 12 min read
Flow Matching vs Diffusion Models: Key Differences and Connections
Data Party THU
Data Party THU
Oct 6, 2025 · Artificial Intelligence

Why Data, Not Architecture, Drives Locality in Diffusion Models

A recent MIT‑Toyota study shows that the locality observed in image diffusion models emerges from the statistical structure of training data rather than from architectural biases, and a simple linear denoiser can replicate this behavior, reshaping how we think about model design.

Data StatisticsU-Netdiffusion models
0 likes · 10 min read
Why Data, Not Architecture, Drives Locality in Diffusion Models
Amap Tech
Amap Tech
Oct 2, 2025 · Artificial Intelligence

How FantasyWorld Unifies Video Generation and 3D Geometry for Consistent Virtual Worlds

FantasyWorld introduces a geometry‑enhanced framework that augments a frozen video diffusion model with a trainable geometry branch, enabling simultaneous video representation and implicit 3D field generation, achieving spatially consistent, high‑quality virtual worlds and outperforming recent baselines in multi‑view coherence and geometric fidelity.

3D modelingVideo Generationcomputer vision
0 likes · 11 min read
How FantasyWorld Unifies Video Generation and 3D Geometry for Consistent Virtual Worlds
AI2ML AI to Machine Learning
AI2ML AI to Machine Learning
Sep 30, 2025 · Artificial Intelligence

Dynamic Multimodal Video Generation: Prioritizing Stability and High Quality

The article surveys the evolution of video generation models—from early GANs and DCGAN to diffusion‑based approaches like Stable Diffusion and DiT—highlighting how stability, high quality, massive compute, and multimodal data pipelines are shaping the current and future paths of dynamic multimodal video generation.

Stable DiffusionTransformerVideo Generation
0 likes · 7 min read
Dynamic Multimodal Video Generation: Prioritizing Stability and High Quality
Bighead's Algorithm Notes
Bighead's Algorithm Notes
Sep 27, 2025 · Artificial Intelligence

Weekly Time-Series Paper Digest (Sep 20‑26, 2025)

This digest summarizes three recent arXiv papers that propose novel diffusion‑based generation, a channel‑independent convolution for multivariate forecasting, and a style‑guided diffusion framework, each demonstrating improved realism, coherence, and diversity of synthetic time‑series data through extensive experiments.

DS-DiffusionIConvMMD loss
0 likes · 8 min read
Weekly Time-Series Paper Digest (Sep 20‑26, 2025)
Kuaishou Large Model
Kuaishou Large Model
Sep 24, 2025 · Artificial Intelligence

How Generative Reinforcement Learning is Revolutionizing Real-Time Bidding

The article explains the core challenges of real‑time bidding, reviews Kuaishou's evolution from PID to MPC to reinforcement learning, and introduces generative reinforcement‑learning methods (GAVE and CBD) that combine decision transformers or diffusion models with value‑guided exploration and score‑based RTG, achieving significant offline and online performance gains.

advertising algorithmsdiffusion modelsgenerative reinforcement learning
0 likes · 15 min read
How Generative Reinforcement Learning is Revolutionizing Real-Time Bidding
Bighead's Algorithm Notes
Bighead's Algorithm Notes
Sep 12, 2025 · Artificial Intelligence

AI for Finance: Quantum Asset Clustering, Causal Market Troughs, Multimodal Forecasting, Diffusion SDEs

This article summarizes four recent AI‑driven finance papers: a quantum‑annealing asset clustering algorithm, a causal machine‑learning model for predicting market troughs, a multimodal large‑model approach to financial time‑series forecasting, and a diffusion‑model method for generating stochastic‑differential‑equation sample paths.

Asset ClusteringCausal MLMultimodal Forecasting
0 likes · 7 min read
AI for Finance: Quantum Asset Clustering, Causal Market Troughs, Multimodal Forecasting, Diffusion SDEs
Sohu Smart Platform Tech Team
Sohu Smart Platform Tech Team
Sep 12, 2025 · Artificial Intelligence

How AI is Revolutionizing Video Creation: From Text‑to‑Video to Real‑Time Editing

This article systematically explores the technical evolution, core principles, and emerging innovations of AI‑generated video, covering generation methods, GAN and diffusion models, transformer‑based DiT architectures, efficiency‑boosting NCR, audio‑visual V2A integration, and real‑world applications across media, education, and commerce.

AI video generationGaNNCR
0 likes · 25 min read
How AI is Revolutionizing Video Creation: From Text‑to‑Video to Real‑Time Editing
Bighead's Algorithm Notes
Bighead's Algorithm Notes
Sep 5, 2025 · Artificial Intelligence

Weekly Quantitative Finance Paper Digest (Aug 30 – Sep 5, 2025)

This digest reviews four recent AI‑driven finance papers: a robust MCVaR portfolio optimizer with ellipsoidal support and RKHS uncertainty, a PPO‑based adaptive weighting system for LLM‑generated alphas, an empirical comparison of price‑based, GICS‑based, and LLM‑embedding stock clustering, and a diffusion‑model approach that generates future financial chart images from current charts and text prompts.

Large Language Modelsdiffusion modelsportfolio optimization
0 likes · 9 min read
Weekly Quantitative Finance Paper Digest (Aug 30 – Sep 5, 2025)
Data Party THU
Data Party THU
Sep 3, 2025 · Artificial Intelligence

Exploring Multimodal Generative AI: A Tsinghua Tutorial at IJCAI 2025

This article introduces a 1.5‑hour tutorial presented by Tsinghua researchers at IJCAI 2025, covering the latest advances in multimodal generative AI, including multimodal large language models, diffusion models, post‑training generalization techniques, and unified understanding‑generation frameworks.

Generative ModelsIJCAI 2025Large Language Models
0 likes · 5 min read
Exploring Multimodal Generative AI: A Tsinghua Tutorial at IJCAI 2025
AIWalker
AIWalker
Aug 19, 2025 · Artificial Intelligence

DynamicFace: Controllable High‑Quality Face Swapping for Images and Video

DynamicFace introduces a diffusion‑based framework that explicitly decouples identity, pose, expression, illumination and background using composable 3D facial priors, achieving superior identity preservation, motion consistency and visual fidelity in both image and video face‑swapping tasks.

3D facial priorscontrollable generationdiffusion models
0 likes · 13 min read
DynamicFace: Controllable High‑Quality Face Swapping for Images and Video
Xiaohongshu Tech REDtech
Xiaohongshu Tech REDtech
Aug 19, 2025 · Artificial Intelligence

How Single Trajectory Distillation Boosts Diffusion Model Speed and Style Quality

The paper introduces Single Trajectory Distillation (STD), a novel training framework that aligns full PF‑ODE trajectories from a fixed noisy state, uses a Trajectory Bank to cut training cost, and adds an Asymmetric Adversarial Loss to markedly improve style consistency and aesthetic quality while accelerating image and video style‑transfer diffusion models.

AI accelerationStyle Transferconsistency models
0 likes · 14 min read
How Single Trajectory Distillation Boosts Diffusion Model Speed and Style Quality
Data Party THU
Data Party THU
Aug 15, 2025 · Artificial Intelligence

What’s Next for Visual Reinforcement Learning? A Comprehensive 2024‑2025 Survey

This article provides a critical, up‑to‑date overview of visual reinforcement learning, formalizes the problem, traces policy‑optimization evolution, categorizes over 200 recent works into four pillars, analyzes algorithms, reward design, benchmarks, and highlights open challenges and future research directions.

RLHFdiffusion modelsmultimodal AI
0 likes · 7 min read
What’s Next for Visual Reinforcement Learning? A Comprehensive 2024‑2025 Survey
Baidu Geek Talk
Baidu Geek Talk
Aug 11, 2025 · Artificial Intelligence

FLUX-Lightning Slashes Diffusion Inference to 4 Steps, Doubling Speed

FLUX-Lightning, introduced by PaddleMIX, combines phased consistency distillation, adversarial learning, distribution‑matching distillation, and reflow loss to reduce diffusion model inference to just four steps while preserving image quality, and leverages the CINN compiler to achieve over 30% speed gains on A800 GPUs, surpassing existing SOTA acceleration methods.

AI InferenceCINNFlux
0 likes · 21 min read
FLUX-Lightning Slashes Diffusion Inference to 4 Steps, Doubling Speed
Data Party THU
Data Party THU
Aug 9, 2025 · Artificial Intelligence

How SADA Boosts Diffusion Model Sampling Speed by Up to 1.8× Without Losing Quality

The paper introduces SADA (Stability‑guided Adaptive Diffusion Acceleration), a novel paradigm that dynamically allocates sparsity per token using a unified stability criterion, enabling efficient ODE‑based sampling for diffusion and flow‑matching models, achieving up to 1.8× speedup with negligible fidelity loss across SD‑2, SDXL, Flux, ControlNet and MusicLDM.

ODEdiffusion modelsgenerative AI
0 likes · 5 min read
How SADA Boosts Diffusion Model Sampling Speed by Up to 1.8× Without Losing Quality
AI Frontier Lectures
AI Frontier Lectures
Jul 30, 2025 · Artificial Intelligence

DualReal: Seamless Identity and Motion Customization for Video Generation

DualReal introduces a novel adaptive joint training framework that simultaneously customizes subject identity and motion dynamics in video generation, overcoming the conflicts of traditional isolated approaches by using a dual-domain perception adapter and stage-fusion controller, achieving up to 31.8% improvement on CLIP‑I and DINO‑I metrics.

Video Generationdiffusion modelsdual-domain adaptation
0 likes · 13 min read
DualReal: Seamless Identity and Motion Customization for Video Generation
Alimama Tech
Alimama Tech
Jul 23, 2025 · Artificial Intelligence

How Differentiable Solver Search Accelerates Diffusion Model Sampling

This article presents a differentiable solver search method that quickly finds high‑quality sampling paths for diffusion models, demonstrating significant FID improvements across Rectified‑Flow, DDPM/VP, and text‑to‑image models while requiring no model parameter changes.

AIdifferentiable solverdiffusion models
0 likes · 20 min read
How Differentiable Solver Search Accelerates Diffusion Model Sampling
Kuaishou Tech
Kuaishou Tech
Jul 22, 2025 · Artificial Intelligence

How Orthus Achieves Lossless Multimodal Generation with a Unified Autoregressive Transformer

Orthus, a new unified multimodal model presented at ICML 2025, leverages an autoregressive Transformer backbone with separate language and diffusion heads to enable lossless image‑text interleaved generation, outperforming existing models on both understanding and generation benchmarks while remaining computationally efficient.

AI researchautoregressive transformerdiffusion models
0 likes · 11 min read
How Orthus Achieves Lossless Multimodal Generation with a Unified Autoregressive Transformer
AI Frontier Lectures
AI Frontier Lectures
Jul 13, 2025 · Artificial Intelligence

How HarmoniCa Boosts Diffusion Model Speed with Joint Training‑Inference Caching

HarmoniCa, a new feature‑caching framework co‑designed by HKUST, Beihang University, and SenseTime, tackles diffusion model inference bottlenecks by aligning training and inference through Step‑Wise Denoising Training and an Image Error Proxy Objective, achieving up to 2× speedup while preserving image quality.

Performance Accelerationdiffusion modelsfeature caching
0 likes · 9 min read
How HarmoniCa Boosts Diffusion Model Speed with Joint Training‑Inference Caching
Amap Tech
Amap Tech
Jul 11, 2025 · Artificial Intelligence

Unified Self‑Supervised Pretraining Accelerates Image Generation and Improves Understanding

The USP framework introduces masked latent modeling within a VAE space to pre‑train ViT encoders, enabling seamless weight transfer to both image classification, segmentation, and diffusion‑based generation tasks, dramatically speeding up DiT and SiT models while preserving strong visual representations.

Self-supervised LearningVAEViT
0 likes · 13 min read
Unified Self‑Supervised Pretraining Accelerates Image Generation and Improves Understanding
Amap Tech
Amap Tech
Jul 11, 2025 · Artificial Intelligence

Unified Self‑Supervised Pretraining Boosts Image Generation and Understanding

The USP framework introduces masked latent modeling within a VAE space to pretrain ViT encoders, enabling seamless weight transfer to both image classification and diffusion‑based generation tasks, dramatically accelerating training while preserving strong performance across multiple benchmarks.

Self-supervised LearningVision Transformerdiffusion models
0 likes · 10 min read
Unified Self‑Supervised Pretraining Boosts Image Generation and Understanding
Amap Tech
Amap Tech
Jul 9, 2025 · Artificial Intelligence

VMBench: Perception-Aligned Motion Benchmark & LD‑RPS Zero‑Shot Restoration

This article introduces VMBench, the first perception‑aligned video motion generation benchmark that defines a five‑dimensional metric suite and a meta‑guided prompt generation pipeline, and presents LD‑RPS, a zero‑shot unified image restoration framework based on latent diffusion recurrent posterior sampling, together with extensive experiments validating both systems.

Video Generationbenchmarkdiffusion models
0 likes · 14 min read
VMBench: Perception-Aligned Motion Benchmark & LD‑RPS Zero‑Shot Restoration
Amap Tech
Amap Tech
Jul 9, 2025 · Artificial Intelligence

Bridging Human Perception and Video Motion Generation: VMBench & LD‑RPS

This article introduces VMBench, a perception‑aligned video motion generation benchmark with a five‑dimensional metric suite and meta‑guided prompt generation, and LD‑RPS, a zero‑shot unified image restoration framework using latent diffusion and recurrent posterior sampling, detailing their motivations, innovations, experiments, and future directions.

AI researchVideo Generationdiffusion models
0 likes · 14 min read
Bridging Human Perception and Video Motion Generation: VMBench & LD‑RPS
Kuaishou Large Model
Kuaishou Large Model
Jul 3, 2025 · Artificial Intelligence

How EvoSearch Boosts Image & Video Generation with Test‑Time Evolutionary Search

The EvoSearch method introduced by HKUST and Kuaishou’s KuaLing team leverages test‑time scaling to dramatically improve diffusion‑based image and video generation without training, using evolutionary search along the denoising trajectory, achieving state‑of‑the‑art results on SD2.1, Flux‑1‑dev and other models.

Evolutionary SearchVideo Generationdiffusion models
0 likes · 8 min read
How EvoSearch Boosts Image & Video Generation with Test‑Time Evolutionary Search
Tencent Technical Engineering
Tencent Technical Engineering
Jul 3, 2025 · Artificial Intelligence

Winning the NTIRE 2025 UGC Video Enhancement Challenge: A Progressive AI Framework

Tencent’s TEG team secured first place in the NTIRE 2025 UGC Video Enhancement competition by introducing a progressive, three‑stage AI framework that decomposes enhancement tasks into expert models for color correction, denoising, and temporal stability, incorporates advanced loss functions, extensive hardware‑level optimizations, INT8 quantization techniques, and outlines future diffusion‑based generative enhancements.

AIdiffusion modelshardware optimization
0 likes · 17 min read
Winning the NTIRE 2025 UGC Video Enhancement Challenge: A Progressive AI Framework
Kuaishou Tech
Kuaishou Tech
Jul 2, 2025 · Artificial Intelligence

How EvoSearch Supercharges Image and Video Generation with Test‑Time Evolutionary Search

EvoSearch, a test‑time evolutionary search method, dramatically improves image and video generation by increasing inference compute without extra training, outperforming existing scaling techniques on diffusion and flow models while maintaining robustness and diversity across multiple benchmarks.

AI researchEvolutionary SearchVideo Generation
0 likes · 8 min read
How EvoSearch Supercharges Image and Video Generation with Test‑Time Evolutionary Search
Kuaishou Large Model
Kuaishou Large Model
Jun 11, 2025 · Artificial Intelligence

12 Kuaishou Breakthrough Papers at CVPR 2025: Video Generation, Diffusion & Multimodal AI

CVPR 2025 in Nashville will feature 12 Kuaishou papers spanning large‑scale video datasets, quality assessment, 3D/4D reconstruction, controllable generation, diffusion scaling laws, multimodal simulation, and novel benchmarks, highlighting the company's cutting‑edge contributions to video AI research.

diffusion modelslarge-scale datasets
0 likes · 21 min read
12 Kuaishou Breakthrough Papers at CVPR 2025: Video Generation, Diffusion & Multimodal AI
DataFunTalk
DataFunTalk
Jun 8, 2025 · Artificial Intelligence

Why Autoregressive Video Models Like MAGI-1 May Outperform Diffusion Approaches

The article examines the current dominance of diffusion models in commercial video generation, contrasts them with autoregressive methods, and details how the open‑source MAGI‑1 model combines both paradigms to achieve longer, more controllable video synthesis while addressing scalability and quality challenges.

AI researchAutoregressive ModelsMAGI-1
0 likes · 70 min read
Why Autoregressive Video Models Like MAGI-1 May Outperform Diffusion Approaches
Amap Tech
Amap Tech
Jun 5, 2025 · Artificial Intelligence

How MVPainter Achieves Accurate, High‑Detail 3D Texture Generation with Multi‑View Diffusion

MVPainter introduces a fully open‑source pipeline that generates high‑quality, PBR‑compatible 3D textures from a single reference image and a white model by leveraging multi‑view diffusion, geometric control, and a human‑aligned evaluation framework, dramatically improving texture fidelity, alignment, and detail.

3D texture generationAIPBR
0 likes · 10 min read
How MVPainter Achieves Accurate, High‑Detail 3D Texture Generation with Multi‑View Diffusion
AntTech
AntTech
Jun 4, 2025 · Artificial Intelligence

LLaDA and LLaDA‑V: Large Language Diffusion Models and Their Multimodal Extensions

This article presents the LLaDA series of diffusion‑based large language models, explains how their generative‑modeling principle yields language intelligence comparable to autoregressive models, and details the multimodal LLaDA‑V architecture, training methods, experimental results, and broader implications for AI research.

Generative ModelingLarge Language Modelsdiffusion models
0 likes · 10 min read
LLaDA and LLaDA‑V: Large Language Diffusion Models and Their Multimodal Extensions
AI Frontier Lectures
AI Frontier Lectures
May 23, 2025 · Artificial Intelligence

How SuperEdit Boosts Instruction-Based Image Editing with Rectified Supervision

SuperEdit introduces rectified instruction generation and contrastive supervision to fix noisy supervision in instruction‑based image editing, achieving up to 9.19% performance gains on Real‑Edit benchmarks without extra model parameters or pre‑training, and releases all data and code publicly.

diffusion modelsimage editingvisual language models
0 likes · 15 min read
How SuperEdit Boosts Instruction-Based Image Editing with Rectified Supervision
AIWalker
AIWalker
May 16, 2025 · Artificial Intelligence

GPDiT Sets New SOTA in Video Generation with Faster, Unified Diffusion‑Autoregressive Framework

GPDiT, a novel autoregressive diffusion transformer, unifies diffusion and autoregressive modeling for video generation, introducing lightweight causal attention and a parameter‑free rotation‑based time conditioning that boost temporal consistency and cut training/inference costs, achieving state‑of‑the‑art results on multiple benchmarks.

Video Generationautoregressive modelingcausal attention
0 likes · 16 min read
GPDiT Sets New SOTA in Video Generation with Faster, Unified Diffusion‑Autoregressive Framework
AI Algorithm Path
AI Algorithm Path
May 15, 2025 · Artificial Intelligence

Understanding Diffusion Models: Core Principles Explained

This article explains the fundamental principles of diffusion models, using physics and machine‑learning analogies to describe forward and reverse diffusion, the role of Gaussian noise, iteration trade‑offs, U‑Net architecture, and shared‑weight training for image generation.

U-Netdiffusion modelsforward diffusion
0 likes · 8 min read
Understanding Diffusion Models: Core Principles Explained