Tagged articles

computer vision

687 articles · Page 3 of 7
AntTech
AntTech
Sep 3, 2024 · Artificial Intelligence

2024 Inclusion Bund Conference AI Innovation Competition and Deepfake Challenge Results

The 2024 Inclusion Bund Conference in Shanghai announced the winners of its newly added AI Innovation Competition, including the AFAC Financial Intelligence Contest and the Global Deepfake Attack‑Defense Challenge, highlighting participation from over 7,000 teams across more than 20 countries and showcasing cutting‑edge deepfake detection achievements.

AIFinTechInnovation Competition
0 likes · 7 min read
2024 Inclusion Bund Conference AI Innovation Competition and Deepfake Challenge Results
JD Cloud Developers
JD Cloud Developers
Aug 29, 2024 · Artificial Intelligence

How AI Powers E‑Commerce Content Compliance and Price Governance

This article explains how e‑commerce platforms use AI‑driven content compliance to detect malicious products, price manipulation, and counterfeit goods, outlining the technical challenges, core business metrics, model‑based solutions for price over‑pricing, and personal growth advice for compliance engineers.

AINLPcomputer vision
0 likes · 9 min read
How AI Powers E‑Commerce Content Compliance and Price Governance
Bilibili Tech
Bilibili Tech
Aug 27, 2024 · Artificial Intelligence

Multimodal Video Scene Classification for Adaptive Video Processing

The paper presents a multimodal video scene classification system that leverages CLIP‑generated pseudo‑labels and a fine‑tuned image encoder to automatically identify nature, animation/game, and document scenes, enabling more effective adaptive transcoding, intelligent restoration, and quality assessment for user‑generated content on platforms such as Bilibili.

Bilibili multimediaCLIPResNet
0 likes · 17 min read
Multimodal Video Scene Classification for Adaptive Video Processing
Rare Earth Juejin Tech Community
Rare Earth Juejin Tech Community
Aug 22, 2024 · Artificial Intelligence

Understanding Faster R-CNN: Architecture, Training, and Experimental Results

This article provides an in‑depth overview of the Faster R‑CNN object detection framework, covering its background, key innovations such as the Region Proposal Network, detailed algorithmic principles, training procedures, experimental results on PASCAL VOC and MS COCO, and a reproducible PyTorch implementation.

Faster R-CNNPyTorchRPN
0 likes · 14 min read
Understanding Faster R-CNN: Architecture, Training, and Experimental Results
php Courses
php Courses
Jul 26, 2024 · Artificial Intelligence

Real-Time Image Processing with PHP and OpenCV

This tutorial explains how PHP developers can install OpenCV and the php-opencv extension, write code to capture webcam video, display live frames in a browser, and perform real-time face detection using computer‑vision techniques.

PHPcomputer visionimage-processing
0 likes · 6 min read
Real-Time Image Processing with PHP and OpenCV
Baidu Geek Talk
Baidu Geek Talk
Jul 24, 2024 · Artificial Intelligence

AI-Driven Fusion of Peking Opera Characters with Ink-Wash Painting Style Using PaddleGAN

Li Yilin’s AI project blends Peking Opera characters with traditional ink‑wash painting by using PaddleHub for style transfer and PaddleGAN’s First‑Order Motion model for facial motion, then adds music and Wav2Lip lip‑sync, producing videos that modernize Chinese heritage and gauge public cultural awareness.

AIPaddleGANPeking Opera
0 likes · 9 min read
AI-Driven Fusion of Peking Opera Characters with Ink-Wash Painting Style Using PaddleGAN
Full-Stack Cultivation Path
Full-Stack Cultivation Path
Jul 17, 2024 · Artificial Intelligence

Open-Source PDF Toolkit Delivers High-Accuracy Layout and Formula Detection

PDF‑Extract‑Kit is an open‑source toolkit that combines high‑accuracy layout detection, formula detection, formula recognition, and OCR for PDFs, and the article details its model comparisons, evaluation on academic and textbook datasets, and step‑by‑step instructions for running it on Windows or macOS, including Apple Silicon.

OCROpen SourcePDF-Extract-Kit
0 likes · 6 min read
Open-Source PDF Toolkit Delivers High-Accuracy Layout and Formula Detection
Kuaishou Tech
Kuaishou Tech
Jul 16, 2024 · Artificial Intelligence

LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control

LivePortrait is an open‑source, controllable portrait video generation framework that transfers facial expressions and poses from a driving video to static or dynamic portraits in real time, leveraging a 69M‑frame mixed video‑image training set, stitching and retargeting modules, and achieving high quality with low latency.

AIVideo Animationcomputer vision
0 likes · 14 min read
LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Jul 15, 2024 · Artificial Intelligence

How EasyAnimate v3 Generates High‑Resolution Videos with Diffusion Transformers

EasyAnimate v3, an open‑source video generation system from Alibaba Cloud AI Platform, introduces Diffusion Transformer‑based architecture, Hybrid Motion Module, and Slice VAE to enable image‑to‑video, text‑to‑video, and unlimited‑length video creation with up to 720p/144 fps resolution on modest GPU memory.

AIEasyAnimateVideo Generation
0 likes · 5 min read
How EasyAnimate v3 Generates High‑Resolution Videos with Diffusion Transformers
Selected Java Interview Questions
Selected Java Interview Questions
Jul 3, 2024 · Artificial Intelligence

Integrating OpenCV with Java and Spring Boot for Face Detection and Recognition

This guide provides a comprehensive walkthrough of installing OpenCV, using its Java API for image and video face detection, implementing face comparison, creating custom GUI windows, and integrating the library into a Spring Boot application with detailed code examples and common troubleshooting tips.

Custom GUIOpenCVSpring Boot
0 likes · 25 min read
Integrating OpenCV with Java and Spring Boot for Face Detection and Recognition
Kuaishou Tech
Kuaishou Tech
Jul 1, 2024 · Artificial Intelligence

Short-Form Video Quality Assessment Competition at CVPR NTIRE 2024: Dataset, Challenge Overview, and Top Winning Solutions

The CVPR NTIRE 2024 short-form video quality assessment competition introduced the KVQ dataset, attracted over 200 teams, evaluated submissions using SROCC and PLCC metrics, and highlighted the winning approaches of SJTU MMLab, IH‑VQA, and TVQE, showcasing advances in AI‑driven video quality evaluation.

AI competitionNTIRE 2024computer vision
0 likes · 9 min read
Short-Form Video Quality Assessment Competition at CVPR NTIRE 2024: Dataset, Challenge Overview, and Top Winning Solutions
DaTaobao Tech
DaTaobao Tech
Jul 1, 2024 · Artificial Intelligence

Recent Progress in Vision-Language Models (VLMs)

Over the past year, Vision‑Language Models have surged from early multimodal experiments to competitive open‑source systems rivaling GPT‑4, driven by higher‑resolution processing, richer vision encoders, better projection layers, and larger curated datasets, yet they still face evaluation difficulties, hallucinations, speed limits, and limited multimodal output.

Large Language ModelsVision-Language Modelscomputer vision
0 likes · 24 min read
Recent Progress in Vision-Language Models (VLMs)
Kuaishou Large Model
Kuaishou Large Model
Jun 27, 2024 · Artificial Intelligence

How I2V-Adapter Turns Images into Videos with Minimal Training

Fast‑forwarding image‑to‑video generation, the article introduces I2V‑Adapter, a lightweight plug‑in for Stable Diffusion‑based video diffusion models that converts a single static image into a coherent video without altering the original T2V architecture, and details its design, frame‑similarity prior, experimental results, and real‑world applications.

AII2V-AdapterStable Diffusion
0 likes · 9 min read
How I2V-Adapter Turns Images into Videos with Minimal Training
Kuaishou Tech
Kuaishou Tech
Jun 26, 2024 · Artificial Intelligence

I2V-Adapter: A Lightweight Image‑to‑Video Adapter for Stable Diffusion Video Diffusion Models

The I2V-Adapter paper introduces a plug‑and‑play lightweight module that enables static images to be converted into dynamic videos using Stable Diffusion‑based text‑to‑video diffusion models without altering the original architecture or pretrained parameters, achieving competitive quality with far less training cost.

AII2V-AdapterStable Diffusion
0 likes · 8 min read
I2V-Adapter: A Lightweight Image‑to‑Video Adapter for Stable Diffusion Video Diffusion Models
Ops Development & AI Practice
Ops Development & AI Practice
Jun 22, 2024 · Artificial Intelligence

Why Transformers Revolutionized AI: From NLP to Vision and Speech

Transformers, introduced in 2017, have reshaped neural networks by leveraging attention mechanisms to outperform RNNs and CNNs across NLP, computer vision, and speech tasks, offering parallel processing, long‑range dependency capture, and versatile applications such as translation, text generation, image classification, and speech recognition.

Attention MechanismNLPTransformer
0 likes · 6 min read
Why Transformers Revolutionized AI: From NLP to Vision and Speech
AntTech
AntTech
Jun 18, 2024 · Artificial Intelligence

Ant Group’s 24 Papers Featured at CVPR2024: Topics and Abstracts

The IEEE CVPR2024 conference in Seattle accepted 2,719 papers out of 11,532 submissions, and Ant Group contributed 24 papers covering computer vision, deep learning, digital humans, large models, multimodal remote sensing, vision‑language distillation, federated incremental learning, model‑stealing defense, and more, with one highlighted as a highlight.

Ant GroupCVPR2024Multimodal
0 likes · 17 min read
Ant Group’s 24 Papers Featured at CVPR2024: Topics and Abstracts
Model Perspective
Model Perspective
Jun 17, 2024 · Artificial Intelligence

Can Diffusion Equations Restore Damaged Paintings? A Practical Guide

This article explains how diffusion equation methods can be applied to digitally repair spotted paintings, covering the mathematical representation of images, the underlying heat‑transfer analogy, step‑by‑step inpainting procedures, and improvements such as total variation flow to preserve edges.

Mathematicscomputer visiondiffusion equation
0 likes · 4 min read
Can Diffusion Equations Restore Damaged Paintings? A Practical Guide
Rare Earth Juejin Tech Community
Rare Earth Juejin Tech Community
Jun 16, 2024 · Artificial Intelligence

HRNet Source Code Walkthrough: Keypoint Dataset Construction, Online Data Augmentation, and Training Pipeline

This article provides a detailed, English-language walkthrough of the HRNet source code, covering how the COCO keypoint dataset is built, the online data‑augmentation techniques applied during training, and the end‑to‑end training and inference procedures for human pose estimation.

Data AugmentationHRNetPyTorch
0 likes · 36 min read
HRNet Source Code Walkthrough: Keypoint Dataset Construction, Online Data Augmentation, and Training Pipeline
Meituan Technology Team
Meituan Technology Team
Jun 13, 2024 · Artificial Intelligence

Overview of Meituan's Selected CVPR 2024 Papers and Online Sharing Event

Meituan's tech team highlights seven CVPR 2024 papers—spanning OCR pre‑training, long‑tail semi‑supervised learning, visual AIGC, audio‑visual segmentation and synthetic‑data detection—provides detailed abstracts and experimental results, and announces an online author‑talk session on June 27.

Audio-Visual SegmentationCVPR 2024OCR
0 likes · 18 min read
Overview of Meituan's Selected CVPR 2024 Papers and Online Sharing Event
DataFunSummit
DataFunSummit
Jun 11, 2024 · Artificial Intelligence

AI Technology Evolution, Commercial Drivers, and Practical Applications in Iron Spectrum Image Recognition, Smart Clause Libraries, and Fire Detection

This article examines the commercial forces behind AI technology evolution, explores academic research on iron‑spectrum image recognition, and details product‑level deployments such as smart clause libraries and fire‑detection systems, while highlighting challenges, strategic approaches, and future outlooks for AI adoption.

AIIndustry ApplicationsProduct Development
0 likes · 21 min read
AI Technology Evolution, Commercial Drivers, and Practical Applications in Iron Spectrum Image Recognition, Smart Clause Libraries, and Fire Detection
php Courses
php Courses
May 30, 2024 · Artificial Intelligence

Real-Time Face Recognition with PHP and OpenCV

This article demonstrates how to set up a PHP environment with OpenCV, control a camera to capture images, and implement real-time face detection and recognition using Haar cascades and LBPH algorithms, providing code examples for building a security-oriented facial recognition system.

OpenCVPHPcomputer vision
0 likes · 6 min read
Real-Time Face Recognition with PHP and OpenCV
Liangxu Linux
Liangxu Linux
May 26, 2024 · Artificial Intelligence

Can Palette-Based Recoloring Transform Pokémon Images Without Neural Networks?

This article presents a mathematically modeled algorithm that extracts color palettes from any Pokémon image and applies them to another, optimizing the swap via deep‑feature distance and dense color‑transform space, demonstrating superior visual results and subjective evaluations compared to traditional hue‑shift and other recoloring methods.

color transfercomputer visiondeep learning
0 likes · 14 min read
Can Palette-Based Recoloring Transform Pokémon Images Without Neural Networks?
Huolala Tech
Huolala Tech
May 23, 2024 · Artificial Intelligence

How to Detect and Remove Moiré Patterns with AI and Diffusion Models

This article explains the nature of moiré patterns in digital imaging, reviews manual mitigation techniques, introduces direct and indirect AI‑based recognition methods—including traditional feature extraction and deep‑learning models such as CNNs and diffusion frameworks—and details practical applications and evaluation metrics used by Huolala.

AIcomputer visiondeep learning
0 likes · 17 min read
How to Detect and Remove Moiré Patterns with AI and Diffusion Models
Meituan Technology Team
Meituan Technology Team
May 16, 2024 · Artificial Intelligence

CMIngre: A Cross‑Modal Ingredient‑Level Dataset for Chinese Food Understanding

The CMIngre dataset, created by Meituan’s R&D platform and Tianjin University, offers 8,001 image‑text pairs of 429 Chinese dishes with 95,290 ingredient bounding boxes, enabling fine‑grained ingredient detection and cross‑modal retrieval tasks, and baseline experiments show DINO and CLIP models achieve the strongest performance.

computer visioncross-modal retrievalfood understanding
0 likes · 44 min read
CMIngre: A Cross‑Modal Ingredient‑Level Dataset for Chinese Food Understanding
Xiaohongshu Tech REDtech
Xiaohongshu Tech REDtech
May 10, 2024 · Artificial Intelligence

SIF3D: Sense‑Informed Forecasting of 3D Human Motion with Multimodal Attention

SIF3D is a scene‑aware 3D human motion forecasting framework that fuses observed motion, 3D point‑cloud scenes, and gaze through novel ternary intention‑aware and semantic‑coherence‑aware attention mechanisms, encoding with PointNet++ and Transformers, and decoding with a graph‑convolutional network, achieving state‑of‑the‑art results on GIMO and GTA‑1M benchmarks.

3D scene understandingCVPR2024SIF3D
0 likes · 15 min read
SIF3D: Sense‑Informed Forecasting of 3D Human Motion with Multimodal Attention
AntTech
AntTech
May 6, 2024 · Artificial Intelligence

AGAP: A Simple Technique for Editing 3D Neural Radiance Fields Explained in Plain Language

The article introduces AGAP, an open‑source method from Ant Research Institute that replaces complex 3D neural radiance field editing with a 2D appearance aggregation approach, enabling easy, loss‑less 3D image manipulation comparable to 2D Photoshop, and is showcased in a short explanatory video.

3D editingAGAPNeural Radiance Fields
0 likes · 3 min read
AGAP: A Simple Technique for Editing 3D Neural Radiance Fields Explained in Plain Language
Alimama Tech
Alimama Tech
Apr 24, 2024 · Artificial Intelligence

Mask‑Guided Diffusion for Precise Product Image Generation

Mask‑Guided Diffusion combines instance‑mask training, Masked Canny ControlNet, and Mask‑guided Attribute Binding to preserve product details, correctly bind attributes, fix hand distortion, and generate uniform colored backgrounds, enabling merchants to quickly create high‑quality, controllable product images with Stable Diffusion.

AIControlNetMask Guidance
0 likes · 16 min read
Mask‑Guided Diffusion for Precise Product Image Generation
php Courses
php Courses
Apr 16, 2024 · Artificial Intelligence

Using PHP and OpenCV for Camera‑Based Object Detection

This tutorial explains how to install required libraries, write PHP code that captures images from a webcam, uses OpenCV and php‑facedetect to detect faces, and displays the results with annotated bounding boxes, providing a foundation for further object detection projects.

CameraOpenCVPHP
0 likes · 6 min read
Using PHP and OpenCV for Camera‑Based Object Detection
php Courses
php Courses
Apr 11, 2024 · Artificial Intelligence

How to Use PHP and OpenCV for Real-Time Camera Image Processing

This tutorial explains how PHP developers can install OpenCV and the php‑opencv extension, capture video from a webcam, display live frames in a browser, and perform basic real‑time image processing such as face detection using OpenCV’s cascade classifier.

OpenCVPHPReal-Time Image Processing
0 likes · 5 min read
How to Use PHP and OpenCV for Real-Time Camera Image Processing
Kuaishou Tech
Kuaishou Tech
Mar 6, 2024 · Artificial Intelligence

Short Video Quality Assessment Competition (KVQ) at CVPR NTIRE 2024

The CVPR NTIRE 2024 workshop hosts the first short‑video quality assessment competition, introducing the KVQ dataset of 4,200 videos across nine scenes, providing training/validation data, a baseline 3D Swin‑Transformer model, detailed competition rules, rewards, and organizer contacts.

AICompetitioncomputer vision
0 likes · 7 min read
Short Video Quality Assessment Competition (KVQ) at CVPR NTIRE 2024
DaTaobao Tech
DaTaobao Tech
Mar 6, 2024 · Artificial Intelligence

AI Clothing Graffiti Project: Implementation and Optimization of AIGC Technology in Taobao Life 2

The AI Clothing Graffiti Project in Taobao Life 2 leverages Stable Diffusion, ControlNet, and LoRA to let users generate and stylize clothing designs via text‑image prompts, employing parallel processing, face repair, and content filtering, and has launched successfully, inviting algorithm engineers to join the team.

AIAIGCControlNet
0 likes · 14 min read
AI Clothing Graffiti Project: Implementation and Optimization of AIGC Technology in Taobao Life 2
DataFunSummit
DataFunSummit
Mar 5, 2024 · Artificial Intelligence

AI-Driven Intelligent Management and Regulation of Mold Temperature in Smart Manufacturing

This article explores how artificial intelligence, computer vision, and control algorithms are applied to smart manufacturing for intelligent mold temperature detection, cooling flow regulation, and full‑process system alerts, presenting a detailed solution architecture, key technologies, and a real‑world case study.

AIMold Temperature ControlPID control
0 likes · 11 min read
AI-Driven Intelligent Management and Regulation of Mold Temperature in Smart Manufacturing
AIWalker
AIWalker
Mar 4, 2024 · Artificial Intelligence

How LLaMA Invades Computer Vision: Introducing VisionLLaMA from Meituan and Zhejiang University

VisionLLaMA extends the LLaMA transformer architecture to 2‑D visual data with naive and pyramid designs, introduces a 2‑D rotary position encoding (AS2DRoPE), and achieves faster convergence and superior performance over state‑of‑the‑art ViT models on image generation, classification, segmentation, and detection tasks.

AS2DRoPELlamaSwiGLU
0 likes · 6 min read
How LLaMA Invades Computer Vision: Introducing VisionLLaMA from Meituan and Zhejiang University
NewBeeNLP
NewBeeNLP
Mar 4, 2024 · Artificial Intelligence

A Curated Tour of Mamba Papers: 25 Cutting‑Edge State‑Space Model Innovations

This article presents a GitHub‑hosted collection of 25 recent research papers on Mamba and its variants, summarizing each work’s core contributions across sequence modeling, vision, medical imaging, graph analysis, and multimodal tasks, and highlighting their performance gains over prior methods.

MambaSequence Modelingcomputer vision
0 likes · 13 min read
A Curated Tour of Mamba Papers: 25 Cutting‑Edge State‑Space Model Innovations
Caiyun Tech Team
Caiyun Tech Team
Feb 22, 2024 · Artificial Intelligence

How to Estimate Visibility Using Atmospheric Tower Photos

This article explains a step‑by‑step method for measuring atmospheric visibility by annotating landmarks in fixed‑angle tower photos, extracting their pixel regions, scoring recognizability with a standard‑deviation‑based metric, and computing the farthest visible object using Python and OpenCV.

OpenCVPythonatmospheric tower
0 likes · 8 min read
How to Estimate Visibility Using Atmospheric Tower Photos
Architects' Tech Alliance
Architects' Tech Alliance
Feb 18, 2024 · Artificial Intelligence

How OpenAI’s Sora Redefines Video Generation with 3‑D Consistency and World Simulation

OpenAI’s Sora model introduces a diffusion‑transformer approach that generates high‑fidelity, 60‑second videos with consistent 3‑D camera motion, long‑term object persistence, and the ability to simulate interactive digital worlds, backed by a detailed technical report and research paper.

Artificial IntelligenceOpenAISora
0 likes · 9 min read
How OpenAI’s Sora Redefines Video Generation with 3‑D Consistency and World Simulation
Xiaohongshu Tech REDtech
Xiaohongshu Tech REDtech
Jan 31, 2024 · Artificial Intelligence

Encoding‑Alignment‑Interaction (EAI) Framework for Full‑Body Human Motion Forecasting

The Encoding‑Alignment‑Interaction (EAI) framework predicts full‑body human motion—including detailed hand joints—by extracting spatio‑temporal features with DCT and GCNs, aligning heterogeneous body‑hand representations via Cross‑Context Alignment, and modeling semantic and physical interactions through Cross‑Context Interaction, achieving state‑of‑the‑art accuracy on the GRAB dataset.

EAI frameworkcomputer visioncross-context alignment
0 likes · 15 min read
Encoding‑Alignment‑Interaction (EAI) Framework for Full‑Body Human Motion Forecasting
DaTaobao Tech
DaTaobao Tech
Jan 31, 2024 · Artificial Intelligence

Highlights of Recent AI Research Papers from Top Conferences (2023)

The article curates standout AI papers from 2023 CCF‑A conferences—including CVPR, ICLR, ACM MM, and INFORMS—showcasing advances such as Swin‑Transformer video quality assessment, cross‑modal e‑commerce product search, transformer‑based vehicle routing heuristics, diffusion‑driven dance generation, and reinforcement‑learning inventory replenishment.

AIcomputer visionmultimedia
0 likes · 23 min read
Highlights of Recent AI Research Papers from Top Conferences (2023)
Huolala Tech
Huolala Tech
Jan 25, 2024 · Artificial Intelligence

How Open‑Vocabulary Detection and Segment‑Anything Are Revolutionizing Visual AI at Huolala

This article reviews traditional computer‑vision tasks—classification, detection, and segmentation—highlights their limitations, introduces open‑vocabulary detection and segment‑anything models such as GLIP, Grounding DINO, and SAM, and details how Huolala applies these advances to driver‑license, packing, and vehicle‑sticker inspections for safer, more efficient AI‑driven operations.

Segmentationcomputer visionobject detection
0 likes · 20 min read
How Open‑Vocabulary Detection and Segment‑Anything Are Revolutionizing Visual AI at Huolala
AIWalker
AIWalker
Jan 22, 2024 · Artificial Intelligence

Depth Anything: An Open-Source Large-Scale Model for Arbitrary Image Depth Estimation

Depth Anything introduces a highly practical monocular depth estimation model that leverages a 62‑million‑image unlabeled dataset, teacher‑student training, strong data perturbations, and DINOv2‑based semantic supervision to achieve zero‑shot capability and state‑of‑the‑art performance over MiDaS across multiple benchmarks.

DINOv2Depth EstimationMonocular Depth
0 likes · 8 min read
Depth Anything: An Open-Source Large-Scale Model for Arbitrary Image Depth Estimation
AsiaInfo Technology: New Tech Exploration
AsiaInfo Technology: New Tech Exploration
Jan 12, 2024 · Artificial Intelligence

Exploring NeRF: From Theory to Real-World 3D Reconstruction Tools

This article introduces Neural Radiance Fields (NeRF) as a cutting‑edge AI technique for high‑quality 3D reconstruction, explains its core principles and advantages, outlines a step‑by‑step building workflow, reviews popular open‑source libraries such as Luma AI, NVIDIA Instant NeRF and NeRFStudio, and offers a forward‑looking summary of its potential and challenges.

3D ReconstructionAINeRF
0 likes · 12 min read
Exploring NeRF: From Theory to Real-World 3D Reconstruction Tools
21CTO
21CTO
Dec 17, 2023 · Artificial Intelligence

Remembering Tang Xiaoyu: The Visionary Behind Modern Facial Recognition

AI pioneer Tang Xiaoyu, co‑founder of SenseTime and former director of leading computer‑vision labs, passed away in December 2023, leaving a legacy of groundbreaking facial‑recognition algorithms, influential mentorship, and a profound impact on the global artificial‑intelligence community.

AI PioneerArtificial Intelligencecomputer vision
0 likes · 7 min read
Remembering Tang Xiaoyu: The Visionary Behind Modern Facial Recognition
We-Design
We-Design
Dec 13, 2023 · Artificial Intelligence

How AI-Powered Beauty Filters Evolved: From Classic Portraits to Real-Time Video Effects

This article traces the evolution of beauty filter technology from ancient artistic enhancements to modern AI-driven real-time video effects, detailing key techniques like face detection, skin smoothing, AR integration, and shifting user preferences, while reflecting on its cultural impact on social media aesthetics.

AIARbeauty filters
0 likes · 9 min read
How AI-Powered Beauty Filters Evolved: From Classic Portraits to Real-Time Video Effects
Airbnb Technology Team
Airbnb Technology Team
Dec 8, 2023 · Artificial Intelligence

Leveraging Image Aesthetics and Photo Sorting Algorithms to Enhance Airbnb Listings

Airbnb’s new computer‑vision pipeline trains a deep‑learning aesthetic model with an EMD loss to rank photos, automatically sorts new‑listing images by design and room type, and scales real‑time similarity search via HNSW‑based ANN on AWS OpenSearch, boosting click‑through, bookings, and enabling unsupervised visual recommendations.

AirbnbEmbedding SearchPhoto Sorting Algorithm
0 likes · 9 min read
Leveraging Image Aesthetics and Photo Sorting Algorithms to Enhance Airbnb Listings
IT Services Circle
IT Services Circle
Dec 6, 2023 · Artificial Intelligence

AI Image Outpainting: Unexpected Transformations and How It Works

The article showcases a series of humorous and surprising AI‑generated image expansions from Douyin, explains the underlying outpainting technology, and discusses why such tools are both entertaining and useful despite occasional odd results.

AIGenerative FillImage Expansion
0 likes · 6 min read
AI Image Outpainting: Unexpected Transformations and How It Works
Python Programming Learning Circle
Python Programming Learning Circle
Nov 30, 2023 · Artificial Intelligence

Common Python Libraries for Computer Vision Projects

This article introduces ten popular Python libraries for computer vision, describing their main features, typical applications, and providing concise code examples to help beginners and practitioners quickly choose and use the right tools for image processing and deep learning tasks.

LibrariesPythoncomputer vision
0 likes · 10 min read
Common Python Libraries for Computer Vision Projects
DataFunTalk
DataFunTalk
Nov 24, 2023 · Artificial Intelligence

Open Vocabulary Detection Contest 2023: Summary of Winning Teams' Technical Solutions

The article reviews the Open Vocabulary Detection Contest organized by the Chinese Society of Image and Graphics and 360 AI Institute, describing the competition setup, dataset characteristics, and detailed winning approaches that combine Detic, CLIP, prompt learning, and multi‑stage pipelines to achieve strong few‑shot and zero‑shot object detection performance.

CLIPCompetitionOpen-Vocabulary Detection
0 likes · 17 min read
Open Vocabulary Detection Contest 2023: Summary of Winning Teams' Technical Solutions
Test Development Learning Exchange
Test Development Learning Exchange
Nov 16, 2023 · Artificial Intelligence

Building a Python Image Editing Tool with Pillow, OpenCV, and NumPy

This guide demonstrates how to create a custom image editing tool in Python by leveraging the Pillow, OpenCV, and NumPy libraries, providing step‑by‑step code examples for opening, resizing, filtering, converting to grayscale, edge detection, rotation, channel manipulation, blurring, contour extraction, and color adjustment.

NumPyPythoncomputer vision
0 likes · 6 min read
Building a Python Image Editing Tool with Pillow, OpenCV, and NumPy
Network Intelligence Research Center (NIRC)
Network Intelligence Research Center (NIRC)
Nov 9, 2023 · Artificial Intelligence

How Wav2Lip Achieves Accurate Speech‑Driven Lip Sync with Expert Discriminators

The article analyzes the limitations of traditional speech‑driven lip‑sync methods and explains how Wav2Lip introduces a pretrained multi‑frame expert sync discriminator, a two‑stage GAN training pipeline, and a specialized generator architecture to produce high‑quality, audio‑aligned facial videos.

GaNWav2Lipaudio‑visual synchronization
0 likes · 7 min read
How Wav2Lip Achieves Accurate Speech‑Driven Lip Sync with Expert Discriminators
Tencent Tech
Tencent Tech
Nov 9, 2023 · Artificial Intelligence

How Adaptive Skinning Model Boosts Low-Cost High-Quality 3D Face Reconstruction

This article introduces the Adaptive Skinning Model (ASM), a low‑cost yet high‑precision 3D face reconstruction technique that leverages Gaussian‑Mixture skinning weights and dynamic bone binding to surpass traditional 3DMM methods and achieve state‑of‑the‑art results on multiple benchmarks.

3D face reconstructionGaussian Mixture Modeladaptive skinning
0 likes · 13 min read
How Adaptive Skinning Model Boosts Low-Cost High-Quality 3D Face Reconstruction
Huawei Cloud Developer Alliance
Huawei Cloud Developer Alliance
Oct 31, 2023 · Artificial Intelligence

Edge‑Cloud AI Powers Student Fatigue‑Driving Detection – Challenge Cup Winners

The 18th Challenge Cup showcased cutting‑edge student projects on fatigue‑driving detection, with Huawei Cloud’s edge‑cloud collaborative topic drawing nearly a thousand participants and five top teams demonstrating AI‑driven solutions that combine incremental training, low‑light enhancement, and lightweight models for real‑time safety alerts.

AIEdge computingFatigue Detection
0 likes · 6 min read
Edge‑Cloud AI Powers Student Fatigue‑Driving Detection – Challenge Cup Winners
Python Programming Learning Circle
Python Programming Learning Circle
Oct 26, 2023 · Artificial Intelligence

Animal Recognition Techniques Using Deep Learning and Image Processing

This article reviews animal recognition technology, covering its background, basic principles, image‑processing, feature extraction, machine‑learning and deep‑learning methods, dataset construction, preprocessing, and feature‑selection techniques, and provides Python code examples for implementing CNNs and traditional classifiers.

animal recognitioncomputer visiondeep learning
0 likes · 18 min read
Animal Recognition Techniques Using Deep Learning and Image Processing
Network Intelligence Research Center (NIRC)
Network Intelligence Research Center (NIRC)
Oct 23, 2023 · Artificial Intelligence

How Multiple‑Instance Learning Boosts Context Understanding in Video Anomaly Detection

The article reviews the CVPR 2021 MIST framework, explaining how a multiple‑instance pseudo‑label generator and a self‑guided attention encoder work together with sparse continuous sampling to improve context awareness and detection accuracy in weakly‑supervised video anomaly detection.

Attention EncoderMultiple Instance LearningSelf‑Training
0 likes · 9 min read
How Multiple‑Instance Learning Boosts Context Understanding in Video Anomaly Detection
DaTaobao Tech
DaTaobao Tech
Oct 13, 2023 · Artificial Intelligence

Understanding Stable Diffusion: Core Principles and Technical Architecture

The article demystifies Stable Diffusion by explaining its low‑cost latent‑space design and conditioning mechanisms, comparing it to autoregressive, VAE, flow‑based and GAN models, detailing the iterative noise‑to‑image process, token‑based text‑to‑image control, version differences, common generation issues, and providing implementation code examples.

AI image generationStable DiffusionVAE
0 likes · 15 min read
Understanding Stable Diffusion: Core Principles and Technical Architecture
Meituan Technology Team
Meituan Technology Team
Oct 11, 2023 · Artificial Intelligence

Meituan Vision AI Research Highlights and Open‑Source Releases

This article compiles Meituan's cutting‑edge computer‑vision research and engineering achievements—including CVPR award‑winning segmentation, YOLOv6 releases, GPU inference optimizations, the Food2K dataset, and numerous paper digests—to provide practical insights for visual AI practitioners.

CVPRFood2KGPU inference
0 likes · 11 min read
Meituan Vision AI Research Highlights and Open‑Source Releases
Kuaishou Large Model
Kuaishou Large Model
Sep 27, 2023 · Artificial Intelligence

DVIS: Decoupled Framework that Sets New SOTA in Video Instance Segmentation

DVIS introduces a decoupled video instance segmentation framework that splits the task into segmentation, tracking, and refinement modules, achieving state-of-the-art performance across VIS, VPS, and VSS benchmarks while maintaining low computational overhead, and demonstrates robustness in both online and offline settings.

Transformercomputer visiondeep learning
0 likes · 12 min read
DVIS: Decoupled Framework that Sets New SOTA in Video Instance Segmentation
Rare Earth Juejin Tech Community
Rare Earth Juejin Tech Community
Sep 16, 2023 · Artificial Intelligence

Understanding DeepSort: A Classic Multi-Object Tracking Algorithm

This article introduces the fundamentals of object tracking in computer vision, explains classic algorithms such as SORT and its deep learning extension DeepSort, describes their underlying mechanisms including Kalman filtering, Hungarian assignment, feature extraction via CNNs, and provides references and code resources for further study.

CNNDeepSortHungarian algorithm
0 likes · 10 min read
Understanding DeepSort: A Classic Multi-Object Tracking Algorithm
Rare Earth Juejin Tech Community
Rare Earth Juejin Tech Community
Aug 26, 2023 · Artificial Intelligence

Using AI and RPA to Solve Slider Captcha: A Practical Implementation with YOLOv8 and PyAutoGUI

This article demonstrates how to combine AI‑based object detection (YOLOv8) with robotic process automation (pyautogui) to automatically locate, drag and release slider captchas, covering data preparation, model training, screen capture, coordinate extraction, mouse simulation, and robustness improvements.

AIRPAYOLOv8
0 likes · 15 min read
Using AI and RPA to Solve Slider Captcha: A Practical Implementation with YOLOv8 and PyAutoGUI
DataFunSummit
DataFunSummit
Aug 24, 2023 · Artificial Intelligence

Panoramic Indoor Layout Estimation with Vision Transformer (PanoViT)

This article introduces the PanoViT model, a vision‑transformer‑based approach for indoor layout estimation from panoramic images, covering its research background, architectural components, experimental results on public datasets, and step‑by‑step usage within ModelScope.

3D ReconstructionModelScopecomputer vision
0 likes · 8 min read
Panoramic Indoor Layout Estimation with Vision Transformer (PanoViT)
Rare Earth Juejin Tech Community
Rare Earth Juejin Tech Community
Aug 24, 2023 · Artificial Intelligence

Neural Style Transfer with PyTorch: Theory and Implementation

This article introduces neural style transfer, explains its underlying principles using VGG19 feature extraction, content and style loss definitions, and provides a complete PyTorch implementation with code for loading images, extracting features, computing Gram matrices, and optimizing the output image.

PyTorchStyle Transfercomputer vision
0 likes · 14 min read
Neural Style Transfer with PyTorch: Theory and Implementation
Top Architect
Top Architect
Aug 22, 2023 · Artificial Intelligence

Face Recognition Search: Principles, Implementation Steps, and Applications

This article explains the background, core principles, preprocessing, feature extraction, matching algorithms, and practical application scenarios of face recognition search, and provides detailed reference implementations with Java and OpenCV code examples for building a complete system.

OpenCVcomputer visiondeep learning
0 likes · 15 min read
Face Recognition Search: Principles, Implementation Steps, and Applications
DaTaobao Tech
DaTaobao Tech
Aug 21, 2023 · Artificial Intelligence

Action Sensitivity Learning for Temporal Action Localization

The paper presents Action Sensitivity Learning (ASL), a framework that models frame‑wise importance at both class‑level (via learnable Gaussian distributions) and instance‑level (using quality scores), integrates these weights into classification and regression losses, adds a contrastive InfoNCE term, and achieves state‑of‑the‑art temporal action localization performance across six benchmark datasets.

Action Sensitivity LearningTemporal Action LocalizationVideo Understanding
0 likes · 8 min read
Action Sensitivity Learning for Temporal Action Localization
Rare Earth Juejin Tech Community
Rare Earth Juejin Tech Community
Aug 17, 2023 · Artificial Intelligence

Getting Started with YOLOv8 on the Ultralytics Platform: Installation, Command‑Line Usage, and Model Training

This article introduces the YOLOv8 object‑detection framework on the Ultralytics platform, covering environment setup, command‑line and Python APIs for inference, model‑file options, result interpretation, data annotation, training procedures, and exporting models to various deployment formats.

PythonUltralyticsYOLO
0 likes · 14 min read
Getting Started with YOLOv8 on the Ultralytics Platform: Installation, Command‑Line Usage, and Model Training
Rare Earth Juejin Tech Community
Rare Earth Juejin Tech Community
Aug 16, 2023 · Artificial Intelligence

Deep Dive into OCR – Chapter 2: Development and Classification of OCR Technology

This article provides a comprehensive overview of OCR technology, detailing the evolution from traditional hand‑crafted methods to modern deep‑learning approaches, describing image preprocessing, text detection and recognition pipelines, summarizing classic machine‑learning algorithms, and presenting a practical OpenCV implementation with Python code.

OCROpenCVPython
0 likes · 23 min read
Deep Dive into OCR – Chapter 2: Development and Classification of OCR Technology
Rare Earth Juejin Tech Community
Rare Earth Juejin Tech Community
Aug 12, 2023 · Artificial Intelligence

An Introduction to OCR: Concepts, History, Applications, Datasets, and Technical Workflow

This article provides a comprehensive overview of Optical Character Recognition (OCR), covering its definition, historical development, classification, real‑world applications, technical pipeline, common challenges, mitigation strategies, popular datasets, model performance comparisons, and leading open‑source platforms.

OCROptical Character Recognitioncomputer vision
0 likes · 16 min read
An Introduction to OCR: Concepts, History, Applications, Datasets, and Technical Workflow
Model Perspective
Model Perspective
Aug 2, 2023 · Artificial Intelligence

How Segment Anything (SAM) Is Revolutionizing Image Segmentation

This article explains the fundamentals of image segmentation, introduces the open‑source Segment Anything Model (SAM) and its massive SA‑1B dataset, outlines SAM's unique promptable, real‑time capabilities, and explores its wide‑ranging future applications across AR/VR, content creation, and scientific research.

AISAMcomputer vision
0 likes · 7 min read
How Segment Anything (SAM) Is Revolutionizing Image Segmentation
Meituan Technology Team
Meituan Technology Team
Jul 27, 2023 · Artificial Intelligence

Street Scene Understanding: Segmentation Technology, Research Progress, and Business Applications

Meituan’s Street‑Scene Understanding team built a high‑precision, efficient segmentation system that aligns motion and static semantics, mines hard examples, iterates models via a data‑model loop, and pursues unified open‑world segmentation, winning multiple CVPR 2023 awards and powering map production, autonomous delivery and store‑scene reconstruction.

AICVPR 2023Meituan
0 likes · 31 min read
Street Scene Understanding: Segmentation Technology, Research Progress, and Business Applications
php Courses
php Courses
Jul 24, 2023 · Artificial Intelligence

Image Edge Enhancement Using PHP and OpenCV

This article explains how to perform image edge enhancement by installing PHP and the OpenCV library, importing images, invoking OpenCV functions, selecting edge detection algorithms such as Sobel or Canny, processing the image with custom code, and displaying or saving the enhanced result.

Edge DetectionOpenCVPHP
0 likes · 5 min read
Image Edge Enhancement Using PHP and OpenCV
Rare Earth Juejin Tech Community
Rare Earth Juejin Tech Community
Jul 24, 2023 · Artificial Intelligence

Understanding Slide-Transformer: An Efficient Local Attention Module for Vision Transformers

This article explains the Slide-Transformer paper, describing how the proposed Slide Attention replaces inefficient Im2Col‑based local attention with depthwise convolutions and a deformable shift module, achieving high efficiency, flexibility, and hardware‑agnostic performance for Vision Transformers.

Deformable ShiftDepthwise ConvolutionSlide Attention
0 likes · 13 min read
Understanding Slide-Transformer: An Efficient Local Attention Module for Vision Transformers
Huolala Tech
Huolala Tech
Jul 21, 2023 · Artificial Intelligence

Visual Language Models Power Open-Set Detection and Surgical Tool Segmentation

Recent advances in visual language models enable zero-shot multimodal tasks, and this article explores their application to open-set object detection, prompt learning, and promptable surgical instrument segmentation, highlighting methods like CLIP, CoOp, and the DetPro framework with experimental results across multiple benchmarks.

MultimodalSemantic Segmentationcomputer vision
0 likes · 12 min read
Visual Language Models Power Open-Set Detection and Surgical Tool Segmentation
php Courses
php Courses
Jul 21, 2023 · Artificial Intelligence

Image Segmentation with PHP and OpenCV

This tutorial explains how to perform image segmentation using the OpenCV library in PHP, covering environment setup, library import, image loading, grayscale conversion, thresholding, result display, and saving the segmented output.

OpenCVPHPcomputer vision
0 likes · 4 min read
Image Segmentation with PHP and OpenCV
php Courses
php Courses
Jul 18, 2023 · Artificial Intelligence

Implementing Face Recognition with PHP and OpenCV

This article provides a step‑by‑step tutorial on installing OpenCV and the PHP OpenCV extension on Ubuntu, then demonstrates how to write PHP code for face detection and recognition using OpenCV's cascade classifier and FisherFaceRecognizer, complete with example scripts and usage instructions.

OpenCVPHPcomputer vision
0 likes · 7 min read
Implementing Face Recognition with PHP and OpenCV
php Courses
php Courses
Jul 17, 2023 · Artificial Intelligence

Implementing Facial Landmark Detection with PHP and OpenCV

This tutorial demonstrates how to set up PHP and OpenCV, install necessary libraries, write and run a PHP script that detects faces and extracts facial landmarks, and saves the annotated image, providing a practical introduction to facial landmark detection in computer vision.

Facial Landmark DetectionOpenCVPHP
0 likes · 5 min read
Implementing Facial Landmark Detection with PHP and OpenCV
Rare Earth Juejin Tech Community
Rare Earth Juejin Tech Community
Jul 12, 2023 · Artificial Intelligence

Comprehensive Guide to Vision Transformer (ViT): Architecture, Patch Tokenization, Embedding, Fine‑tuning, and Performance

This article provides an in‑depth, English‑language overview of Vision Transformer (ViT), covering its Transformer‑based architecture, patch‑to‑token conversion, token and position embeddings, fine‑tuning strategies such as 2‑D interpolation, experimental results versus CNNs, and the model’s broader significance for multimodal AI research.

Fine‑tuningPatch EmbeddingTransformer
0 likes · 25 min read
Comprehensive Guide to Vision Transformer (ViT): Architecture, Patch Tokenization, Embedding, Fine‑tuning, and Performance
Kuaishou Large Model
Kuaishou Large Model
Jul 7, 2023 · Artificial Intelligence

How HairStep Revolutionizes Single-View 3D Hair Reconstruction

This paper introduces HairStep, a novel intermediate representation combining Strand Maps and Depth Maps, and demonstrates how it reduces domain gap and improves single‑view 3D hair reconstruction accuracy across multiple algorithms, supported by new annotated datasets (HiSa, HiDa) and fair evaluation metrics.

3D hair reconstructionHairStepcomputer vision
0 likes · 11 min read
How HairStep Revolutionizes Single-View 3D Hair Reconstruction
Efficient Ops
Efficient Ops
Jun 26, 2023 · Artificial Intelligence

How Multimodal AI Is Revolutionizing Credit Card Fraud Detection

Amid tightening financial regulations, ICBC's software team proposes a multimodal AI anti‑fraud framework that combines image, video, and structured data to detect deep‑fake, mask, and forged‑document attacks, enriches verification with cross‑modal cues, and outlines future expansion to text and speech modalities.

AIMultimodalcomputer vision
0 likes · 7 min read
How Multimodal AI Is Revolutionizing Credit Card Fraud Detection
Programmer DD
Programmer DD
Jun 20, 2023 · Artificial Intelligence

Yann LeCun: Today's AI Still Below Dog Level – Inside Meta’s Voicebox, MusicGen & I‑JEPA

Meta’s chief AI scientist Yann LeCun warned that current large language models still fall short of human and even dog intelligence, citing their lack of real‑world understanding, while Meta unveiled three new generative AI models—Voicebox for speech, MusicGen for music, and I‑JEPA for image reasoning—showcasing both progress and remaining limitations.

Artificial IntelligenceLarge Language Modelscomputer vision
0 likes · 7 min read
Yann LeCun: Today's AI Still Below Dog Level – Inside Meta’s Voicebox, MusicGen & I‑JEPA
Xiaohongshu Tech REDtech
Xiaohongshu Tech REDtech
Jun 20, 2023 · Artificial Intelligence

Open-Vocabulary Object Attribute Recognition with OvarNet: A Unified Framework for Detection and Attribute Classification

At CVPR 2023 the Xiaohongshu team presented OvarNet, a unified one‑stage Faster‑RCNN model built on CLIP that uses prompt learning and knowledge distillation to jointly detect objects and recognize open‑vocabulary attributes, achieving state‑of‑the‑art results on VAW, MS‑COCO, LSA and OVAD datasets.

attribute recognitioncomputer visionknowledge distillation
0 likes · 12 min read
Open-Vocabulary Object Attribute Recognition with OvarNet: A Unified Framework for Detection and Attribute Classification
Meituan Technology Team
Meituan Technology Team
Jun 15, 2023 · Artificial Intelligence

Meituan Technical Team's 8 CVPR 2023 Papers: Overview and Insights

This article reviews eight CVPR 2023 papers selected by Meituan’s technology team, covering self‑supervised learning, domain adaptation, federated learning, object detection, 3D reconstruction, GAN‑based pre‑training, RGB‑T tracking, vision‑language navigation, and visual‑textual layout generation, highlighting each work’s methodology, experiments, and reported performance gains.

3D Object DetectionCVPR 2023Domain Adaptation
0 likes · 15 min read
Meituan Technical Team's 8 CVPR 2023 Papers: Overview and Insights
Alimama Tech
Alimama Tech
Jun 14, 2023 · Artificial Intelligence

Intelligent Live‑Streaming Video Editing Techniques and Practices

Alibaba Mama’s end‑to‑end intelligent clipping system automatically transforms long live‑stream e‑commerce videos into short, high‑quality ads by segmenting streams, classifying speech with GPT‑based tags, selecting visually appealing clips, arranging coherent storylines, and applying effects, achieving 96% classification accuracy and improved advertising efficiency.

AIContent Optimizationcomputer vision
0 likes · 14 min read
Intelligent Live‑Streaming Video Editing Techniques and Practices
Network Intelligence Research Center (NIRC)
Network Intelligence Research Center (NIRC)
Jun 9, 2023 · Artificial Intelligence

2023 NIRC PhD Graduates Reveal Cutting-Edge AI and Network Intelligence Research

In 2023 the Network Intelligent Research Center celebrated its largest PhD graduating class—seven scholars whose dissertations span deep‑vision hand‑gesture estimation, multi‑scenario network transmission, graph alignment, interactive streaming, knowledge‑defined networking, wireless body‑area networking, and more—showcasing significant AI‑driven advances and high‑impact publications.

Artificial IntelligenceGraph AlignmentNetwork Intelligence
0 likes · 30 min read
2023 NIRC PhD Graduates Reveal Cutting-Edge AI and Network Intelligence Research
DataFunSummit
DataFunSummit
May 31, 2023 · Artificial Intelligence

Evolution of Face Detection Techniques: Datasets, Research Directions, and Future Work

This article reviews the evolution of face detection, covering the Widely‑Face dataset, major research directions such as feature fusion, label assignment, auxiliary supervision, anchor‑free methods, NAS‑based designs, summarizes key papers from S3FD to MogFace, introduces ModelScope implementations, and outlines future challenges and opportunities.

AI researchModel Evaluationcomputer vision
0 likes · 13 min read
Evolution of Face Detection Techniques: Datasets, Research Directions, and Future Work
Test Development Learning Exchange
Test Development Learning Exchange
May 27, 2023 · Artificial Intelligence

Eight Essential OpenCV Examples for Image Processing

This article introduces eight fundamental OpenCV examples—including image reading, display, grayscale conversion, edge detection, resizing, Gaussian blur, and face detection—providing concise Python code snippets and explanations to help readers quickly apply these common computer‑vision techniques.

OpenCVPythonTutorial
0 likes · 5 min read
Eight Essential OpenCV Examples for Image Processing
DataFunTalk
DataFunTalk
May 13, 2023 · Artificial Intelligence

Multimedia Content Understanding at Weibo: Video Summarization, Quality Assessment, OCR, Embedding, and CV‑CUDA Optimization

This article presents Weibo's comprehensive multimedia content understanding pipeline, covering video summarization techniques, quality assessment models, OCR advancements, video embedding strategies, and the performance benefits of CV‑CUDA acceleration, while highlighting real‑world applications and engineering trade‑offs.

CV-CUDAEmbeddingOCR
0 likes · 32 min read
Multimedia Content Understanding at Weibo: Video Summarization, Quality Assessment, OCR, Embedding, and CV‑CUDA Optimization
AntTech
AntTech
May 6, 2023 · Artificial Intelligence

Wu Wenjun AI Science and Technology Award Honors Tsinghua and Ant Group's Unconstrained Human Portrait Perception and Understanding Technology

The 2022 Wu Wenjun Artificial Intelligence Science and Technology Award recognized a decade‑long collaborative effort by Tsinghua University and Ant Group's security lab for breakthrough research on unconstrained human portrait perception and understanding, highlighting three core scientific discoveries, extensive academic impact, and large‑scale commercial applications in identity verification.

AI AwardsAnt GroupHuman Portrait Recognition
0 likes · 5 min read
Wu Wenjun AI Science and Technology Award Honors Tsinghua and Ant Group's Unconstrained Human Portrait Perception and Understanding Technology
Baidu Tech Salon
Baidu Tech Salon
Apr 25, 2023 · Game Development

How to Build Sensor‑Free Motion Games with PP‑TinyPose and FastDeploy

This article explains how to develop sensor‑less motion-controlled games by leveraging the PP‑TinyPose keypoint detection model and FastDeploy inference tool, detailing the required setup, code snippets, and a reusable PyQt5 framework for creating webcam‑driven interactive demos.

AIFastDeployPP-TinyPose
0 likes · 11 min read
How to Build Sensor‑Free Motion Games with PP‑TinyPose and FastDeploy
DataFunSummit
DataFunSummit
Apr 20, 2023 · Artificial Intelligence

SenseTime Unveils Multimodal ‘SenseNova’ Large Model System and Its Industry Applications

SenseTime introduced its visual‑centric multimodal large‑model platform SenseNova, detailing model scaling, extensive AI infrastructure, diverse industry deployments such as autonomous driving and generative content, and the challenges of compute efficiency and data acquisition in the race for advanced AI.

AI infrastructurecomputer visionlarge models
0 likes · 13 min read
SenseTime Unveils Multimodal ‘SenseNova’ Large Model System and Its Industry Applications
Baidu Tech Salon
Baidu Tech Salon
Apr 14, 2023 · Artificial Intelligence

How PaddleDepth and Paddle3D Enable Low‑Cost 3D Vision Development

This article examines the challenges of 3D vision data acquisition and explains how Baidu's PaddleDepth and Paddle3D toolkits provide low‑cost depth collection, super‑resolution, and end‑to‑end perception pipelines, showcasing performance on KITTI and Middlebury datasets with code examples.

3D visionDepth EstimationOpen Source
0 likes · 12 min read
How PaddleDepth and Paddle3D Enable Low‑Cost 3D Vision Development
AntTech
AntTech
Apr 12, 2023 · Artificial Intelligence

Ant Technology Research Institute Interactive Intelligence Lab – 13 Papers Accepted at CVPR 2023 and Recent AI Research Highlights

The Ant Technology Research Institute’s Interactive Intelligence Lab announced that 13 of its papers were accepted at CVPR 2023, alongside other recent achievements in generative models and 3D vision, highlighting collaborations with top universities and summarizing the lab’s contributions to artificial intelligence research.

3D visionCVPRGenerative Models
0 likes · 6 min read
Ant Technology Research Institute Interactive Intelligence Lab – 13 Papers Accepted at CVPR 2023 and Recent AI Research Highlights
Baidu Tech Salon
Baidu Tech Salon
Apr 7, 2023 · Artificial Intelligence

Ambiguity-Resistant Semi-supervised Learning (ARSL) for Single-stage Object Detection

ARSL, an ambiguity‑resistant semi‑supervised learning framework for single‑stage object detection, introduces Joint‑Confidence Estimation and Task‑Separation Assignment to resolve selection and assignment ambiguities in pseudo‑labels, thereby markedly improving pseudo‑label quality and achieving state‑of‑the‑art AP gains on COCO benchmarks.

ARSLcomputer visionjoint confidence estimation
0 likes · 8 min read
Ambiguity-Resistant Semi-supervised Learning (ARSL) for Single-stage Object Detection
Baidu Geek Talk
Baidu Geek Talk
Mar 16, 2023 · Artificial Intelligence

PaddleDetection v2.6 Release: PP-YOLOE Family Expansion and Advanced Detection Algorithms

PaddleDetection v2.6 expands the PP‑YOLOE family with rotating, small‑object, dense‑object, and ultra‑lightweight edge‑GPU models, upgrades PP‑Human and PP‑Vehicle toolboxes, releases semi‑supervised, few‑shot and distillation learning methods, adds numerous state‑of‑the‑art algorithms, and improves infrastructure with Python 3.10, EMA filtering and AdamW support.

BaiduPP-YOLOEPaddleDetection
0 likes · 14 min read
PaddleDetection v2.6 Release: PP-YOLOE Family Expansion and Advanced Detection Algorithms
政采云技术
政采云技术
Mar 9, 2023 · Artificial Intelligence

Comprehensive Overview of Object Detection: From Traditional Methods to Modern Deep Learning Models

This article provides a comprehensive overview of object detection, describing traditional sliding‑window approaches, deep‑learning based two‑stage and one‑stage models such as R‑CNN, Faster R‑CNN, YOLO series, and discusses current challenges, improvement directions, and future research trends in the field.

R-CNNYOLOcomputer vision
0 likes · 29 min read
Comprehensive Overview of Object Detection: From Traditional Methods to Modern Deep Learning Models