Tagged articles

computer vision

687 articles · Page 5 of 7
Java Captain
Java Captain
Dec 4, 2021 · Artificial Intelligence

Java Spring Boot License Plate Recognition and Training System (Open‑Source)

This open‑source project implements a Spring Boot and Maven based license‑plate detection and training system in Java, leveraging OpenCV and JavaCPP, supporting multiple plate colors, SVM and ANN algorithms, and providing a B/S architecture with SQLite, Swagger documentation, and extensible image‑recognition features.

OpenCVSpring Bootcomputer vision
0 likes · 4 min read
Java Spring Boot License Plate Recognition and Training System (Open‑Source)
Kuaishou Large Model
Kuaishou Large Model
Dec 3, 2021 · Artificial Intelligence

How Can Your Face Reveal Heart Rate? Exploring rPPG Technology

This article explains the principles of remote photoplethysmography (rPPG), how facial skin color changes caused by heartbeats can be captured by a camera to measure heart rate, respiration, SpO₂ and other physiological signals, and reviews traditional and data‑driven algorithms for robust signal extraction.

AIcomputer visionheart rate detection
0 likes · 7 min read
How Can Your Face Reveal Heart Rate? Exploring rPPG Technology
Kuaishou Tech
Kuaishou Tech
Dec 1, 2021 · Industry Insights

Turning Sketches into Live AR Characters: Kuaishou’s All‑Things‑AR Technical Journey

This article details how Kuaishou transformed a user‑drawn sketch concept into the All‑Things‑AR feature, covering background inspiration, the end‑to‑end pipeline, data collection, mobile‑friendly segmentation model design, model optimizations, engineering integration, SLAM‑based camera localization, and final production results.

ARIndustry Case StudySLAM
0 likes · 15 min read
Turning Sketches into Live AR Characters: Kuaishou’s All‑Things‑AR Technical Journey
21CTO
21CTO
Nov 27, 2021 · Artificial Intelligence

How Huawei’s “Genius Teen” Scaled AutoML to Millions of Phones

Huawei’s 201‑million‑yuan “genius teen” Zhong Zhao leveraged AutoML to deploy high‑precision image‑pixel processing algorithms across tens of millions of Mate and P series smartphones, pioneering large‑scale commercial use of AutoML and advancing mobile visual models with dynamic convolution kernels and adversarial data augmentation.

AutoMLHuaweicomputer vision
0 likes · 9 min read
How Huawei’s “Genius Teen” Scaled AutoML to Millions of Phones
DeWu Technology
DeWu Technology
Nov 18, 2021 · Artificial Intelligence

Background Complexity Detection for Sneaker Images Using MobileNet, FPN, and Modified SAM

The project presents a lightweight MobileNet‑FPN architecture enhanced with a modified spatial‑attention module that evaluates corner‑based self‑similarity to classify sneaker photo backgrounds, achieving 96% test accuracy—exceeding baseline CNN performance—and meeting business targets of over 80% hint accuracy and 90% mandatory enforcement.

CNNMobileNetattention
0 likes · 12 min read
Background Complexity Detection for Sneaker Images Using MobileNet, FPN, and Modified SAM
DataFunTalk
DataFunTalk
Nov 16, 2021 · Artificial Intelligence

InsightFace: Open‑Source 2D/3D Deep Face Analysis Toolbox with PaddlePaddle Support

InsightFace is an open‑source 2D/3D deep face analysis toolbox that implements a variety of detection, alignment and recognition algorithms, now supports PaddlePaddle with out‑of‑the‑box models, high‑throughput distributed training up to 60 million classes, and provides a one‑line demo script for quick testing.

ArcFaceInsightfacePaddlePaddle
0 likes · 3 min read
InsightFace: Open‑Source 2D/3D Deep Face Analysis Toolbox with PaddlePaddle Support
Alibaba Terminal Technology
Alibaba Terminal Technology
Nov 15, 2021 · Artificial Intelligence

How AI Powers Smart Home Workouts on Mobile: Alibaba Sports’ Pose‑Tracking

Alibaba Sports’ AI-powered smart workout system transforms a simple smartphone and a few square meters of space into an interactive home fitness solution, using MNN‑based pose estimation to recognize and correct dozens of exercises, while addressing challenges like accuracy, performance, and automated testing.

AIautomated testingcomputer vision
0 likes · 11 min read
How AI Powers Smart Home Workouts on Mobile: Alibaba Sports’ Pose‑Tracking
Amap Tech
Amap Tech
Nov 4, 2021 · Artificial Intelligence

POI Signboard Image Retrieval: Technical Solution, Model Design, and Future Directions

To efficiently filter unchanged POI signboards, the authors propose a multimodal image‑retrieval system that combines enhanced global and local visual features with BERT‑encoded OCR text, using metric learning and alignment techniques to achieve over 95 % accuracy while handling occlusion, viewpoint variation, and subtle text changes.

POIcomputer visiondeep learning
0 likes · 17 min read
POI Signboard Image Retrieval: Technical Solution, Model Design, and Future Directions
Laiye Technology Team
Laiye Technology Team
Sep 24, 2021 · Artificial Intelligence

Self‑Supervised Learning and Contrastive Methods for Computer Vision and OCR Applications

This article surveys self‑supervised learning techniques for computer‑vision tasks, explains common pretext tasks and contrastive loss designs, reviews representative models such as SimCLR, MoCo, SmAV and SimSiam, and demonstrates their practical impact on a captcha‑OCR system with measurable accuracy gains.

OCRSelf-supervised LearningSimCLR
0 likes · 23 min read
Self‑Supervised Learning and Contrastive Methods for Computer Vision and OCR Applications
Kuaishou Tech
Kuaishou Tech
Sep 17, 2021 · Artificial Intelligence

SnowflakeNet: Point Cloud Completion by Snowflake Point Deconvolution with Skip-Transformer

SnowflakeNet introduces a novel Snowflake Point Deconvolution architecture combined with a Skip-Transformer to progressively split seed points, enabling high‑quality point‑cloud completion that preserves fine‑grained geometric details such as smooth surfaces, sharp edges, and corners across dense and sparse datasets.

3D ReconstructionSnowflakeNetcomputer vision
0 likes · 10 min read
SnowflakeNet: Point Cloud Completion by Snowflake Point Deconvolution with Skip-Transformer
JD Retail Technology
JD Retail Technology
Sep 8, 2021 · Artificial Intelligence

ARShoe: Real-Time Augmented Reality Shoe Try-On System on Smartphones

The paper presents ARShoe, the first practical real‑time augmented reality shoe try‑on system for smartphones, detailing its multi‑branch neural network, foot pose estimation, rendering pipeline, a newly built foot dataset, and extensive experiments demonstrating high accuracy and over 30 FPS performance on multiple devices.

ARaugmented realitycomputer vision
0 likes · 6 min read
ARShoe: Real-Time Augmented Reality Shoe Try-On System on Smartphones
Baidu Geek Talk
Baidu Geek Talk
Sep 8, 2021 · Artificial Intelligence

How PP‑OCRv2 Boosts OCR Speed and Accuracy with Five Key Innovations

The article provides a comprehensive technical overview of PaddleOCR's PP‑OCRv2, detailing its five major algorithmic enhancements, performance improvements over previous versions, historical milestones, core capabilities, and links to the open‑source repositories for developers interested in state‑of‑the‑art OCR solutions.

Data AugmentationModel OptimizationOCR
0 likes · 10 min read
How PP‑OCRv2 Boosts OCR Speed and Accuracy with Five Key Innovations
NetEase Smart Enterprise Tech+
NetEase Smart Enterprise Tech+
Sep 2, 2021 · Artificial Intelligence

How AI Detects Video Deepfakes: Techniques, Challenges, and Real-World Solutions

This article explores the rapid rise of AI‑generated video deepfakes, examines the four main manipulation techniques, discusses the inherent security risks, and presents NetEase Yidun’s comprehensive detection framework—including face‑detection‑based classification, semi‑supervised learning, feature fusion, and model distillation—to combat content‑security threats.

AI securitycomputer visiondeepfake detection
0 likes · 12 min read
How AI Detects Video Deepfakes: Techniques, Challenges, and Real-World Solutions
Kuaishou Large Model
Kuaishou Large Model
Aug 30, 2021 · Artificial Intelligence

How Kuaishou’s Y‑Tech Fixes Background Distortion in Portrait Beautification

This article explains the challenges of background distortion caused by portrait beautification effects, describes Kuaishou Y‑Tech’s line‑segment‑based optimization framework that preserves line slopes and triangle shapes, and demonstrates the algorithm’s effectiveness through before‑and‑after visual results.

Optimizationbackground correctioncomputer vision
0 likes · 11 min read
How Kuaishou’s Y‑Tech Fixes Background Distortion in Portrait Beautification
Beike Product & Technology
Beike Product & Technology
Aug 13, 2021 · Artificial Intelligence

AI-Powered Intelligent Testing Platform for Frontend UI Quality Assurance

The article describes how an AI-driven testing platform combines computer‑vision, OCR, and machine‑learning techniques to automatically detect frontend UI and backend‑related quality issues in mobile apps, outlines its architecture, core capabilities, deployment workflow, and reports successful real‑world deployments and future plans.

AI testingcomputer visionfrontend quality
0 likes · 11 min read
AI-Powered Intelligent Testing Platform for Frontend UI Quality Assurance
Alimama Tech
Alimama Tech
Aug 11, 2021 · Artificial Intelligence

Dynamic Descriptive Model: A Scalable Paradigm for High‑Quality Native Creative Generation

The Dynamic Descriptive Model (DDM) introduces a scalable pipeline that automatically harvests product assets, perceives their visual attributes, encodes designers’ expertise in an extended SVG‑based descriptive language, and generates high‑quality, native‑looking ad creatives at massive scale, delivering 5‑80 % CTR gains and tens of millions of daily outputs.

AIAdvertisingDynamic Descriptive Model
0 likes · 13 min read
Dynamic Descriptive Model: A Scalable Paradigm for High‑Quality Native Creative Generation
MaGe Linux Operations
MaGe Linux Operations
Aug 9, 2021 · Artificial Intelligence

Top Python Libraries for Image Processing: A Practical Guide with Code

This article introduces the most popular Python image‑processing libraries, explains their core features, and provides ready‑to‑run code examples for tasks such as filtering, segmentation, and computer‑vision applications, helping readers quickly start working with images in Python.

NumPycomputer visionimage processing
0 likes · 9 min read
Top Python Libraries for Image Processing: A Practical Guide with Code
iQIYI Technical Product Team
iQIYI Technical Product Team
Aug 6, 2021 · Artificial Intelligence

I2UV-HandNet: High‑Fidelity 3D Hand Mesh Reconstruction from Monocular RGB Images

I2UV-HandNet reconstructs high-fidelity 3D hand meshes from a single RGB image using an AffineNet encoder‑decoder to predict coarse UV maps and an SRNet super‑resolution module, trained on the SuperHandScan dataset, achieving real‑time performance and state‑of‑the‑art benchmark results, and targeting integration into next‑generation VR headsets without external controllers.

3D meshUV mappingVR
0 likes · 11 min read
I2UV-HandNet: High‑Fidelity 3D Hand Mesh Reconstruction from Monocular RGB Images
MaGe Linux Operations
MaGe Linux Operations
Jul 29, 2021 · Artificial Intelligence

Unlock Powerful Face Recognition with Python’s face_recognition Library

This article introduces the open‑source Python library face_recognition, explains how to install it, locate and extract faces, generate 128‑dimensional embeddings, compare faces, detect facial landmarks, apply virtual makeup, and build a simple custom face‑recognition application with complete code examples and visual results.

computer visionface recognitionfacial landmarks
0 likes · 11 min read
Unlock Powerful Face Recognition with Python’s face_recognition Library
Python Programming Learning Circle
Python Programming Learning Circle
Jul 27, 2021 · Artificial Intelligence

Common Python Libraries for Image Processing: Overview and Code Examples

This article introduces the most widely used Python image‑processing libraries—including scikit‑image, NumPy, SciPy, Pillow, OpenCV‑Python, SimpleCV, Mahotas, SimpleITK, pgmagick, and Pycairo—explaining their key features and providing concise code snippets that demonstrate filtering, segmentation, enhancement, and computer‑vision tasks.

NumPyOpenCVcomputer vision
0 likes · 8 min read
Common Python Libraries for Image Processing: Overview and Code Examples
Test Development Learning Exchange
Test Development Learning Exchange
Jul 21, 2021 · Artificial Intelligence

Drawing Shapes on Images with OpenCV in Python

This tutorial demonstrates how to use OpenCV in Python to read an image and draw basic shapes such as rectangles and circles by specifying coordinates, dimensions, colors, and line thickness, then display the edited image in a window.

Drawing ShapesOpenCVPython
0 likes · 2 min read
Drawing Shapes on Images with OpenCV in Python
Test Development Learning Exchange
Test Development Learning Exchange
Jul 20, 2021 · Artificial Intelligence

Resizing Images with Python and OpenCV

This article demonstrates how to use Python's OpenCV library to read an image, display its original dimensions, resize it to a specified size, save the resized image, and handle user input to close the display windows.

OpenCVPythonResize
0 likes · 2 min read
Resizing Images with Python and OpenCV
Test Development Learning Exchange
Test Development Learning Exchange
Jul 17, 2021 · Artificial Intelligence

Face Recognition with OpenCV and Python

This tutorial explains the concept of facial recognition, describes how it works, and provides step‑by‑step instructions and code examples for implementing face detection and identification using OpenCV and Python, including installation, basic image handling, and a complete sample script.

OpenCVPythoncomputer vision
0 likes · 4 min read
Face Recognition with OpenCV and Python
Kuaishou Large Model
Kuaishou Large Model
Jul 15, 2021 · Artificial Intelligence

How Kuaishou’s YKit AI SDK Powers Mass‑Production of Viral Effects

The article details Kuaishou Y‑tech's YKit AI SDK architecture, its unified interface, modular design, performance optimizations, and three real‑world case studies that illustrate how the SDK enables large‑scale, high‑quality short‑video effects across diverse devices while addressing challenges of effect variety, performance, and cost.

AI SDKARcomputer vision
0 likes · 14 min read
How Kuaishou’s YKit AI SDK Powers Mass‑Production of Viral Effects
21CTO
21CTO
Jul 14, 2021 · Artificial Intelligence

How a Chinese PhD’s Vision Research Earned a 2‑Million‑Yuan Huawei Offer

The article profiles Liao Minghui, a recent PhD graduate from Huazhong University of Science and Technology whose groundbreaking work in computer‑vision text detection earned him top honors, multiple patents, and a record‑breaking 2.01 million‑yuan annual salary offer from Huawei’s “Genius Youth” program.

Academic AchievementArtificial IntelligenceHuawei Recruitment
0 likes · 7 min read
How a Chinese PhD’s Vision Research Earned a 2‑Million‑Yuan Huawei Offer
Beike Product & Technology
Beike Product & Technology
Jul 8, 2021 · Artificial Intelligence

Raster‑to‑Vector Floorplan Reconstruction (R2V) for Standardized Housing Layouts

This article presents the motivation, definitions, related work, and a detailed R2V (Raster‑to‑Vector) modeling pipeline—including DNN segmentation, integer programming, and vector standardization—used by Beike to standardize diverse floor‑plan images, discusses challenges, and outlines future directions, while also noting recruitment opportunities.

computer visionfloorplaninteger optimization
0 likes · 20 min read
Raster‑to‑Vector Floorplan Reconstruction (R2V) for Standardized Housing Layouts
Youku Technology
Youku Technology
Jul 8, 2021 · Artificial Intelligence

Key Findings from Alibaba Moku Lab at ACM MM 2021

At ACM MM 2021, Alibaba’s Moku Lab presented four cutting‑edge studies: an interactive video inpainting system using user doodles, a decoupled IoU regression model for object detection, a spatio‑temporal distortion‑aware video quality assessment framework, and a multimodal emotional relationship recognition dataset and benchmark.

Video Inpaintingcomputer visionmultimodal emotion recognition
0 likes · 8 min read
Key Findings from Alibaba Moku Lab at ACM MM 2021
Miss Fresh Tech Team
Miss Fresh Tech Team
Jul 8, 2021 · Artificial Intelligence

How AI Powers Smart Vending Cabinets: From RFID to Deep Learning Detection

This article details the evolution of intelligent vending cabinets, comparing RFID, gravity, dynamic and static vision solutions, and explains how deep‑learning models, data pipelines, and system architectures enable high‑accuracy, low‑loss product detection and automated operations in modern unmanned retail.

AISmart Vendingcomputer vision
0 likes · 36 min read
How AI Powers Smart Vending Cabinets: From RFID to Deep Learning Detection
New Oriental Technology
New Oriental Technology
Jul 8, 2021 · Artificial Intelligence

Paper Detection and Perspective Correction Using OpenCV.js

This article introduces OpenCV.js, explains its basic concepts and demonstrates a complete workflow for detecting and correcting paper images in the browser using JavaScript, including matrix handling, resizing, filtering, edge detection, contour analysis, perspective transformation, and discusses challenges such as noise and incomplete edges.

JavaScriptOpenCV.jsPaper Detection
0 likes · 10 min read
Paper Detection and Perspective Correction Using OpenCV.js
TiPaiPai Technical Team
TiPaiPai Technical Team
Jul 2, 2021 · Artificial Intelligence

How ContourNet and CenterNet Revolutionize Text Detection

This article explains the challenges of scene text detection and introduces two state‑of‑the‑art models, ContourNet and CenterNet, detailing their architectural innovations, loss functions, and how they overcome issues like extreme aspect ratios and anchor‑based inefficiencies.

CenterNetContourNetcomputer vision
0 likes · 7 min read
How ContourNet and CenterNet Revolutionize Text Detection
TiPaiPai Technical Team
TiPaiPai Technical Team
Jun 28, 2021 · Artificial Intelligence

How Deep Learning Unwarps Twisted Document Images: DocUNet & DewarpNet Explained

This article reviews two end‑to‑end deep‑learning approaches—DocUNet (CVPR 2018) and DewarpNet (ICCV 2019)—for correcting warped document images, detailing their network architectures, synthetic data generation, loss functions, experimental results, and the remaining challenges in document dewarping.

OCRcomputer visiondeep learning
0 likes · 14 min read
How Deep Learning Unwarps Twisted Document Images: DocUNet & DewarpNet Explained
TAL Education Technology
TAL Education Technology
Jun 24, 2021 · Artificial Intelligence

GoodFuture AI Institute Wins Four International Championships at CVPR 2021 Across Multiple Vision Challenges

GoodFuture AI Institute secured four international titles at CVPR 2021—including Person In Context, UG²+, ETH‑XGaze, and ActivityNet—showcasing world‑class computer‑vision algorithms for human‑object interaction, low‑light face detection, gaze estimation, and active speaker detection, and highlighting their deployment in educational AI solutions.

AI competitionActive Speaker DetectionCVPR
0 likes · 9 min read
GoodFuture AI Institute Wins Four International Championships at CVPR 2021 Across Multiple Vision Challenges
Alibaba Cloud Developer
Alibaba Cloud Developer
Jun 22, 2021 · Artificial Intelligence

Turning Parking Cameras into AI‑Powered Safety Guardians

A Qingdao University student leveraged Intel Xeon SG1 GPU, OpenVINO and Mask R-CNN to transform existing parking‑lot cameras into an intelligent system that counts vehicles, detects pedestrians in blind spots, and issues real‑time safety alerts, showcasing a practical AI solution for child safety in crowded parking areas.

AIIntel XeonMask R-CNN
0 likes · 5 min read
Turning Parking Cameras into AI‑Powered Safety Guardians
Xianyu Technology
Xianyu Technology
Jun 9, 2021 · Artificial Intelligence

Applying Visual AI Techniques for Image Quality and Duplicate Detection in Xianyu Marketplace

By deploying large‑scale visual AI—including a ResNet‑101 classifier, ArcFace‑trained matching features, clustering‑based sub‑category refinement, and product‑level image indexing—Xianyu’s marketplace dramatically improves image quality, removes duplicates, enhances search relevance and feed diversity, and filters non‑compliant content.

computer visiondeep learningduplicate detection
0 likes · 16 min read
Applying Visual AI Techniques for Image Quality and Duplicate Detection in Xianyu Marketplace
Amap Tech
Amap Tech
Jun 4, 2021 · Artificial Intelligence

Deploying Multiple CNN Models on Low‑End Devices with MNN: Memory Tricks and Error Debugging

This article explains how a high‑traffic map service captures road features using client‑side computer‑vision models, details the deployment of many CNNs with the lightweight MNN engine on memory‑constrained devices, and shares practical memory‑saving techniques, inference scheduling, and error‑analysis methods.

AndroidMNNMemory Optimization
0 likes · 12 min read
Deploying Multiple CNN Models on Low‑End Devices with MNN: Memory Tricks and Error Debugging
Meituan Technology Team
Meituan Technology Team
Jun 3, 2021 · Artificial Intelligence

LargeFineFoodAI Workshop and Challenge at ICCV 2021

At ICCV 2021 in Montreal, the LargeFineFoodAI workshop—co‑organized by Meituan Vision Intelligence Center, the Chinese Academy of Sciences, Beijing Zhiyuan and the University of Barcelona—will showcase state‑of‑the‑art fine‑grained food image research, feature invited speakers Jain, Aizawa and Radeva, and host a $12,000 prize challenge on Food2K across recognition and retrieval tracks.

Fine-Grained ClassificationICCV 2021challenge
0 likes · 7 min read
LargeFineFoodAI Workshop and Challenge at ICCV 2021
Meituan Technology Team
Meituan Technology Team
May 27, 2021 · Artificial Intelligence

Standardizing Food Delivery Dish Names: Knowledge Graph Construction and Applications

The paper outlines an end‑to‑end pipeline that standardizes highly personalized food‑delivery dish names by combining rule‑based and BERT‑DSSM text synonym detection with EfficientNet image classification, constructing a multi‑level taxonomy that improves aggregation, supply‑demand analysis, recall ranking and merchant tagging.

Food DeliveryKnowledge GraphNLP
0 likes · 17 min read
Standardizing Food Delivery Dish Names: Knowledge Graph Construction and Applications
Tencent Advertising Technology
Tencent Advertising Technology
May 27, 2021 · Artificial Intelligence

Multimodal Video Ad Second-Level Parsing: Algorithm Design and Baseline Analysis for the 2021 Tencent Advertising Algorithm Competition

This article details the algorithmic framework and baseline models for the 2021 Tencent Advertising Algorithm Competition, focusing on multimodal video ad parsing through temporal localization, scene segmentation, and multi-label classification to enhance advertising effectiveness and creative analysis.

Temporal Segmentationadvertising technologycomputer vision
0 likes · 22 min read
Multimodal Video Ad Second-Level Parsing: Algorithm Design and Baseline Analysis for the 2021 Tencent Advertising Algorithm Competition
Kuaishou Tech
Kuaishou Tech
May 24, 2021 · Artificial Intelligence

BCNet: A Bilayer Instance Segmentation Network for Occlusion‑Aware Object Detection

The paper proposes BCNet, a lightweight bilayer instance segmentation network that explicitly models occluder and occludee relationships by treating each region of interest as two overlapping layers, achieving significant performance gains on COCO, COCOA and KINS datasets under heavy occlusion.

bilayer networkcomputer visiondeep learning
0 likes · 10 min read
BCNet: A Bilayer Instance Segmentation Network for Occlusion‑Aware Object Detection
Alimama Tech
Alimama Tech
May 20, 2021 · Artificial Intelligence

How Alibaba’s AI Powers Brand Risk Detection: Models, Data, and Results

This article details Alibaba's AliMama brand risk identification system, covering the challenges of counterfeit detection, the construction of large‑scale brand datasets, the design of classification, logo detection, and variation models, their optimization, evaluation metrics, and future directions for AI‑driven brand protection.

AIAlibababrand risk detection
0 likes · 22 min read
How Alibaba’s AI Powers Brand Risk Detection: Models, Data, and Results
Kuaishou Large Model
Kuaishou Large Model
May 13, 2021 · Artificial Intelligence

How Regressive Domain Adaptation Boosts Unsupervised Keypoint Detection

This article reviews the CVPR2021 paper on Regressive Domain Adaptation (RegDA) for unsupervised keypoint detection, explaining its motivation, novel adversarial regression framework, sparse output-space modeling, min‑min training strategy, extensive experiments, and the resulting performance gains across multiple datasets.

Domain AdaptationUnsupervised Learningcomputer vision
0 likes · 13 min read
How Regressive Domain Adaptation Boosts Unsupervised Keypoint Detection
Kuaishou Tech
Kuaishou Tech
May 10, 2021 · Artificial Intelligence

Semantic Image Matting: Integrating Alpha Pattern Semantics into the Matting Framework

The article presents Semantic Image Matting, a novel approach that incorporates 20 semantic Alpha pattern categories into the matting pipeline via semantic Trimap, region‑based classifiers, multi‑class discriminators, and learnable gradient loss, achieving state‑of‑the‑art results on multiple benchmarks.

Semantic Segmentationalpha patternscomputer vision
0 likes · 11 min read
Semantic Image Matting: Integrating Alpha Pattern Semantics into the Matting Framework
JD Cloud Developers
JD Cloud Developers
Apr 30, 2021 · Artificial Intelligence

How Face Keypoint Localization Advances Under Masked Conditions: Insights from JD AI’s 3rd Competition

The JD AI Institute and ICME2021 concluded their third face keypoint localization contest, emphasizing efficient masked‑face detection to aid COVID‑19 contact tracing, attracting top universities and tech firms, expanding data scale, and tightening model efficiency constraints to push the field forward.

AI competitioncomputer visiondeep learning
0 likes · 4 min read
How Face Keypoint Localization Advances Under Masked Conditions: Insights from JD AI’s 3rd Competition
Amap Tech
Amap Tech
Apr 23, 2021 · Artificial Intelligence

Design Principles and Implementation of Gaode AR Navigation

The article explains Gaode Maps’ AR navigation design, detailing how environmental factors, spatial experience, color hierarchy, safety considerations, and competitor insights shape a six‑point design framework, and describes prototype testing, implementation strategies for overlapping alerts, and future prospects such as virtual road barriers and multimodal travel.

AR navigationDesign Principlesaugmented reality
0 likes · 8 min read
Design Principles and Implementation of Gaode AR Navigation
360 Quality & Efficiency
360 Quality & Efficiency
Apr 16, 2021 · Artificial Intelligence

Applying YOLOv5 Object Detection for Black, Color, and Blank Screen Classification in Video Frames

This article presents a method that replaces manual visual inspection with an automated YOLOv5‑based object detection pipeline to classify video frames as normal, colorful, or black screens, detailing data annotation, training, loss calculation, inference code, and showing a 97% accuracy improvement over ResNet.

PythonYOLOv5computer vision
0 likes · 11 min read
Applying YOLOv5 Object Detection for Black, Color, and Blank Screen Classification in Video Frames
MaGe Linux Operations
MaGe Linux Operations
Apr 13, 2021 · Artificial Intelligence

Top 10 Free Python Libraries for Image Processing You Should Try

Discover ten essential, free Python libraries for image processing—from scikit-image and NumPy to OpenCV-Python and Pycairo—each with resources, usage examples, and visual demonstrations, enabling you to manipulate, analyze, and transform images efficiently for computer vision and data science projects.

LibrariesOpenCVPython
0 likes · 12 min read
Top 10 Free Python Libraries for Image Processing You Should Try
58UXD
58UXD
Apr 12, 2021 · Artificial Intelligence

How 58.com Built an AI Designer: From Smart Cutout to Intelligent Creative Platform

This article chronicles 58.com’s journey from a small brainstorming room to a full‑scale AI design platform, detailing the development of smart cutout, the BASNet segmentation model, custom loss functions, template editing, and the measurable business impact of the AI designer.

AI designBASNetcomputer vision
0 likes · 15 min read
How 58.com Built an AI Designer: From Smart Cutout to Intelligent Creative Platform
DataFunTalk
DataFunTalk
Apr 10, 2021 · Artificial Intelligence

2020 Computer Vision Breakthroughs: Self‑Supervised Learning, Transformer Attention Modeling, and Neural Radiance Fields

The talk reviews three major 2020 advances in computer vision—self‑supervised learning surpassing supervised pre‑training, the successful adoption of Transformer‑based attention models for detection and classification, and the emergence of Neural Radiance Fields for view synthesis—while highlighting related research from Microsoft Research Asia and the broader community.

2020AI breakthroughsSelf-supervised Learning
0 likes · 19 min read
2020 Computer Vision Breakthroughs: Self‑Supervised Learning, Transformer Attention Modeling, and Neural Radiance Fields
Youku Technology
Youku Technology
Apr 8, 2021 · Artificial Intelligence

Champion Solution of Media AI Alibaba Entertainment Video Object Segmentation Challenge

The Youku AI team won the Media AI Alibaba Entertainment Video Object Segmentation Challenge by enhancing the STM model with a spatial‑constrained memory reader, ASPP‑HRNet refinement, ResNeSt‑101 backbone, and a multi‑stage training pipeline, while also devising an unsupervised framework that combines DetectoRS detection, HRNet mask refinement, STM‑based association, and key‑frame optimization to achieve 95.5% test score on a large, richly annotated video dataset.

Spatial Memory NetworksUnsupervised VOScomputer vision
0 likes · 13 min read
Champion Solution of Media AI Alibaba Entertainment Video Object Segmentation Challenge
Kuaishou Tech
Kuaishou Tech
Apr 6, 2021 · Artificial Intelligence

Frequency-Aware Feature Learning with Single-Center Loss for Face Forgery Detection

Researchers from USTC and Kuaishou propose a frequency‑aware feature learning framework that combines a data‑driven adaptive frequency module with a novel single‑center loss, achieving state‑of‑the‑art performance on deepfake detection while addressing class‑distribution challenges.

AI securitycomputer visiondeepfake detection
0 likes · 7 min read
Frequency-Aware Feature Learning with Single-Center Loss for Face Forgery Detection
Kuaishou Large Model
Kuaishou Large Model
Apr 1, 2021 · Artificial Intelligence

How Kuaishou Y‑Tech Leverages GANs for Real‑Time Face Attribute Editing in Short Videos

This article details Kuaishou Y‑Tech's practical deployment of GAN‑based high‑precision face attribute editing—covering gender, age, hair, and expression transformations—for short‑video effects, discussing background, business applications, technical challenges, and solutions across data preparation, model training, and mobile deployment.

GaNKuaishouStyleGAN
0 likes · 15 min read
How Kuaishou Y‑Tech Leverages GANs for Real‑Time Face Attribute Editing in Short Videos
iQIYI Technical Product Team
iQIYI Technical Product Team
Mar 26, 2021 · Artificial Intelligence

Insights into OCR Technology at iQIYI: Development, Challenges, and Applications

iQIYI’s OCR journey, explained by researcher Harlon, covers the evolution from separate detection and recognition pipelines to end‑to‑end models, key algorithms like CTPN, DB and CRNN, large‑scale simulated training, diverse video‑text applications, and future goals such as mobile deployment and tighter NLP integration.

AIOCRPaddleOCR
0 likes · 21 min read
Insights into OCR Technology at iQIYI: Development, Challenges, and Applications
58 Tech
58 Tech
Mar 24, 2021 · Artificial Intelligence

Automated Detection of Illegal Watermarks in Images Using Deep Learning at 58.com

This article describes how 58.com built an end‑to‑end deep‑learning watermark detection service, covering business needs, data collection and augmentation, model selection and iterative improvements (Faster‑RCNN, SSD, YOLOv3, anchor‑free methods), deployment results, and future research directions.

Image ModerationModel Optimizationcomputer vision
0 likes · 14 min read
Automated Detection of Illegal Watermarks in Images Using Deep Learning at 58.com
Huawei Cloud Developer Alliance
Huawei Cloud Developer Alliance
Mar 23, 2021 · Artificial Intelligence

How to Recognize Credit Card Numbers with OpenCV: A Step‑by‑Step Tutorial

This tutorial walks through a project‑based OpenCV workflow that reads a digit template, preprocesses both template and credit‑card images, extracts individual numbers, matches them against the template, and finally overlays the recognized digits onto the original image, illustrating core computer‑vision techniques.

OCROpenCVPython
0 likes · 10 min read
How to Recognize Credit Card Numbers with OpenCV: A Step‑by‑Step Tutorial
Amap Tech
Amap Tech
Mar 22, 2021 · Artificial Intelligence

Visual Technology for Automated POI Name Generation: STR, Text Detection, and Naming Practices

Amap’s visual‑technology pipeline automatically generates and updates POI names by crowdsourcing street‑level images, applying deep‑learning scene‑text recognition, dual‑branch classification of text attributes, and a BERT‑plus‑graph‑attention model that selects and orders recognized text, achieving about 95 % naming accuracy.

Name GenerationOCRPOI
0 likes · 14 min read
Visual Technology for Automated POI Name Generation: STR, Text Detection, and Naming Practices
JD Cloud Developers
JD Cloud Developers
Mar 8, 2021 · Artificial Intelligence

How AI Voice Synthesis Brings ‘Hi, Mom’ to Life: From Film to Real‑World Tech

The article explores how modern AI technologies such as speech synthesis, natural language understanding, and the FastReID computer‑vision library enable realistic voice recreation and cross‑temporal dialogue, turning the emotional premise of the movie “Hi, Mom” into a tangible technical demonstration.

AIFastReIDNatural Language Understanding
0 likes · 10 min read
How AI Voice Synthesis Brings ‘Hi, Mom’ to Life: From Film to Real‑World Tech
Tencent Cloud Developer
Tencent Cloud Developer
Mar 4, 2021 · Artificial Intelligence

WeChat OCR: Implementation of Image Text Extraction Feature

WeChat’s 8.0 update introduced an OCR pipeline that first quickly detects text in images, classifies the image type, applies a lightweight multi‑language detection network and a MobileNetV3‑based DBNet recognizer with a multi‑task CTC/Attention model, then merges results via a rule‑based layout analyzer to deliver accurate, well‑formatted extracted text across diverse languages and document types.

DBNetOCROptical Character Recognition
0 likes · 13 min read
WeChat OCR: Implementation of Image Text Extraction Feature
Laravel Tech Community
Laravel Tech Community
Feb 28, 2021 · Artificial Intelligence

How the “Ant Ya Hey” AI Effect Works and How to Create It

This article explains the popular Douyin AI effect “Ant Ya Hey”, showcases celebrity demos, provides a step‑by‑step guide using Avatarify and video editors, and delves into the underlying First‑Order Motion Model research that powers the realistic facial animation.

AIAvatarifyDouyin
0 likes · 6 min read
How the “Ant Ya Hey” AI Effect Works and How to Create It
Kuaishou Large Model
Kuaishou Large Model
Feb 25, 2021 · Artificial Intelligence

How Kuaishou’s AI‑Powered Beauty Engine Transforms Real‑Time Video

This article details Kuaishou Y‑tech’s Gorgeous beauty platform, covering traditional smoothing, advanced skin‑tone effects, AI‑driven blemish removal, clarity enhancement, local facial tuning, and the UNet‑based GorgeousGAN that delivers one‑click high‑definition beauty for live‑stream and short‑video applications.

AI beautycomputer visiondeep learning
0 likes · 13 min read
How Kuaishou’s AI‑Powered Beauty Engine Transforms Real‑Time Video
360 Tech Engineering
360 Tech Engineering
Feb 23, 2021 · Artificial Intelligence

Video Stutter Detection via Frame Difference Analysis Using FFmpeg

This article explains a method for detecting video stutter by converting uploaded videos into frame sequences with ffmpeg, calculating pixel differences between consecutive frames, aggregating motion metrics, removing scene‑change effects, computing a dynamic factor, and outputting a binary result indicating the presence or absence of stutter.

Algorithmcomputer visionframe analysis
0 likes · 5 min read
Video Stutter Detection via Frame Difference Analysis Using FFmpeg
JD Cloud Developers
JD Cloud Developers
Feb 10, 2021 · Artificial Intelligence

How JD Tech’s Breakthrough AI Papers Dominated AAAI 2021

JD Tech showcased a remarkable 21-paper presence at AAAI 2021, covering federated learning, spatio‑temporal AI, recommendation systems, computer vision, and causal learning, highlighting the company’s transition from research to real‑world AI applications across smart cities, retail, and risk management.

AAAI 2021Recommendation Systemscausal learning
0 likes · 12 min read
How JD Tech’s Breakthrough AI Papers Dominated AAAI 2021
ByteFE
ByteFE
Feb 9, 2021 · Fundamentals

Curated Self‑Study Resources for Emerging Tech Fields (Multimedia, AI, CV, RL, MT, Knowledge Graph, Mobile, Frontend)

This guide compiles recommended books, courses, and open‑source projects across multimedia, artificial intelligence, computer vision, reinforcement learning, machine translation, knowledge graphs, Android, iOS, and frontend development to help newcomers and job seekers systematically deepen their technical expertise.

Artificial IntelligenceResourcescomputer vision
0 likes · 12 min read
Curated Self‑Study Resources for Emerging Tech Fields (Multimedia, AI, CV, RL, MT, Knowledge Graph, Mobile, Frontend)
iQIYI Technical Product Team
iQIYI Technical Product Team
Feb 5, 2021 · Game Development

AR+AI Powered Video Interactive Mini‑Games on iQIYI: Architecture, Face & Gesture Control, and Lua Game Layer

iQIYI’s AR+AI powered video interactive mini‑games blend a custom VideoAR engine with real‑time AI‑driven face and gesture detection, use lightweight Lua for game logic, and offer rapid hot‑updates, enabling diverse IP integrations that have attracted over a million participants and boosted viewer engagement.

AIARLua
0 likes · 12 min read
AR+AI Powered Video Interactive Mini‑Games on iQIYI: Architecture, Face & Gesture Control, and Lua Game Layer
Amap Tech
Amap Tech
Feb 1, 2021 · Artificial Intelligence

AMAP-TECH Algorithm Competition: Dynamic Road Condition Analysis Using In-Vehicle Video

The AMAP‑TECH competition challenged participants to infer real‑time road conditions from in‑vehicle video, prompting the authors to combine lane‑wise vehicle detection with LightGBM and later an end‑to‑end DenseNet‑GRU model, augment data, ensemble five networks, and achieve a 0.7237 F1 score while outlining future deployment and research directions.

Model deploymentcomputer visiondeep learning
0 likes · 15 min read
AMAP-TECH Algorithm Competition: Dynamic Road Condition Analysis Using In-Vehicle Video
Kuaishou Large Model
Kuaishou Large Model
Jan 28, 2021 · Artificial Intelligence

How Portrait Deformation Powers Modern Beauty Filters: Algorithms Explained

This article explores the core portrait deformation techniques behind today’s beauty and body‑shaping effects—covering affine transforms, Moving Least Squares, triangulation, liquify, offset, 3D mesh, and deep‑learning approaches—detailing their principles, implementations, and visual results in live‑streaming and short‑video apps.

AIbeauty filterscomputer vision
0 likes · 13 min read
How Portrait Deformation Powers Modern Beauty Filters: Algorithms Explained
iQIYI Technical Product Team
iQIYI Technical Product Team
Jan 15, 2021 · Artificial Intelligence

How AI is Transforming Video Creation and Consumption at Scale

The article examines how iQIYI leverages AI across the video ecosystem—from intelligent material search, old‑film restoration, and voice cloning to virtual idols, XR production, and AI‑driven advertising—to boost creator efficiency, enhance user experience, and accelerate industry-wide digital transformation.

AIIndustry InsightsVoice Cloning
0 likes · 14 min read
How AI is Transforming Video Creation and Consumption at Scale
Amap Tech
Amap Tech
Jan 15, 2021 · Artificial Intelligence

Solution Overview of the AMAP-TECH Algorithm Competition: Dynamic Road Condition Analysis from In‑Vehicle Video Images

To tackle the AMAP‑TECH competition’s dynamic road‑condition classification from scarce, imbalanced vehicle‑video frames, the team combined YOLOv5 object detection, ResNeXt101‑based semantic embeddings, and engineered temporal detection statistics, feeding the fused features into a five‑fold LightGBM model that achieved top weighted‑F1 performance.

LightGBMResNeXtYOLOv5
0 likes · 10 min read
Solution Overview of the AMAP-TECH Algorithm Competition: Dynamic Road Condition Analysis from In‑Vehicle Video Images
Didi Tech
Didi Tech
Dec 29, 2020 · Artificial Intelligence

Evolution and Challenges of Perception in L4 Autonomous Driving

The article traces L4 autonomous-driving perception from early rule-based point-cloud methods through data-driven deep-learning models to emerging self-learning, multi-task systems, and highlights four key hurdles—model generalization and explainability, robust multi-sensor fusion, real-time compute limits, and proper uncertainty handling—calling for integrated AI, engineering, and data solutions.

AIcomputer visiondeep learning
0 likes · 12 min read
Evolution and Challenges of Perception in L4 Autonomous Driving
Meituan Technology Team
Meituan Technology Team
Dec 24, 2020 · Artificial Intelligence

Meituan Unmanned Delivery Technical Salon – AI Research on Instance Segmentation, Visual Localization, Trajectory Prediction, and Depth‑Pose Learning

On January 9, 2021, Meituan hosted an unmanned‑delivery technical salon in Beijing where experts presented cutting‑edge AI research—including the CenterMask instance‑segmentation method, 3D geometry‑aware camera localization, multi‑agent trajectory prediction with attention‑based spatio‑temporal graphs, real‑time stereo visual‑inertial odometry calibration, and self‑supervised depth‑pose learning for dynamic scenes.

AISelf-supervised Learningautonomous driving
0 likes · 7 min read
Meituan Unmanned Delivery Technical Salon – AI Research on Instance Segmentation, Visual Localization, Trajectory Prediction, and Depth‑Pose Learning
Suning Technology
Suning Technology
Dec 17, 2020 · Artificial Intelligence

How AI Powers SuNing’s Unmanned Stores: From Face Detection to Smart Retail

This article outlines SuNing's unmanned store technology, comparing its data-driven, product selection, and customer experience advantages over traditional shops, and detailing AI-powered applications such as face detection, target tracking, image recognition, and 3D reconstruction that enable 24‑hour service, intelligent merchandising, and precise customer analytics.

AIData AnalyticsRetail Technology
0 likes · 24 min read
How AI Powers SuNing’s Unmanned Stores: From Face Detection to Smart Retail
DataFunTalk
DataFunTalk
Dec 9, 2020 · Artificial Intelligence

WeChat Identify: From Object Detection to Large‑Scale Image Search – Technical Overview

This article details the evolution of WeChat’s Identify product, explaining its end‑to‑end image recognition pipeline—including object detection, multi‑label classification, mobile‑side detection, large‑scale retrieval, unsupervised clustering, and system architecture—while showcasing various application scenarios such as product, plant, and landmark recognition.

Large-Scale RetrievalWeChatcomputer vision
0 likes · 12 min read
WeChat Identify: From Object Detection to Large‑Scale Image Search – Technical Overview
Python Crawling & Data Mining
Python Crawling & Data Mining
Dec 9, 2020 · Artificial Intelligence

Unlock 3D Human Pose Capture with FrankMocap: A Powerful Open‑Source AI Tool

FrankMocap, an open‑source AI algorithm from Facebook AI Research and HKU, delivers simultaneous 3D full‑body and hand pose estimation from a single monocular video, runs at about 9.5 FPS on a RTX 2080, and includes easy installation steps, code examples, and links to its GitHub repository and paper.

3D pose estimationOpen SourcePython
0 likes · 6 min read
Unlock 3D Human Pose Capture with FrankMocap: A Powerful Open‑Source AI Tool
Top Architect
Top Architect
Dec 4, 2020 · Artificial Intelligence

Java-based ID Card OCR Project Using OpenCV, JavaCPP, and Tess4J

This article introduces a Java OCR project for ID cards that integrates OpenCV, JavaCPP, and Tess4J to perform image preprocessing, region cropping, and character recognition without requiring OpenCV installation, and details its features, encountered issues, system requirements, updates, and source repository.

ID CardJavaCPPOCR
0 likes · 4 min read
Java-based ID Card OCR Project Using OpenCV, JavaCPP, and Tess4J
DataFunSummit
DataFunSummit
Dec 3, 2020 · Artificial Intelligence

GAN Fundamentals, Variants, and Practical Applications in Image Style Transfer and Handwriting Font Generation

This article provides a comprehensive overview of Generative Adversarial Networks, covering their original formulation, training dynamics, loss functions, major variants such as DCGAN and WGAN, and practical implementations for image‑to‑image translation, style transfer, and handwriting font synthesis at Laiye Technology.

GaNGenerative Adversarial NetworksStyle Transfer
0 likes · 28 min read
GAN Fundamentals, Variants, and Practical Applications in Image Style Transfer and Handwriting Font Generation
Kuaishou Large Model
Kuaishou Large Model
Dec 3, 2020 · Artificial Intelligence

Kuaishou Y‑Tech’s Real‑Time, High‑Precision Facial & Body Keypoint Detection Explained

Y‑Tech’s in‑house keypoint detection system powers Kuaishou’s beauty and effect filters across live streaming, video creation, and editing by leveraging lightweight deep‑learning models, extensive multi‑scenario data collection, and specialized handling of occlusion, enabling real‑time, robust facial and body landmark tracking on diverse mobile devices.

beauty filterscomputer visiondeep learning
0 likes · 10 min read
Kuaishou Y‑Tech’s Real‑Time, High‑Precision Facial & Body Keypoint Detection Explained
360 Quality & Efficiency
360 Quality & Efficiency
Nov 27, 2020 · Artificial Intelligence

Image Similarity Detection Methods: Hashing, Histograms, Feature Matching, BOW+K‑Means, and CNN‑Based Approaches

This article reviews common image similarity detection techniques—including hash-based methods (aHash, pHash, dHash), histogram comparison, feature matching with ORB and SIFT/SURF, bag‑of‑words with K‑Means, and CNN‑based VGG16 features—detailing their algorithms, Python implementations, performance characteristics, and practical considerations.

Hashingcomputer visiondeep learning
0 likes · 15 min read
Image Similarity Detection Methods: Hashing, Histograms, Feature Matching, BOW+K‑Means, and CNN‑Based Approaches
Suning Technology
Suning Technology
Nov 26, 2020 · Artificial Intelligence

How Low-Cost AI Powers Full-Scale Store Digitalization

Li Yongxiang, technical director at Suning Tech, outlines how AI-driven visual unmanned stores and integrated big‑data, cloud, and edge computing solutions enable low‑cost digital transformation across thousands of retail outlets, improving shopper experience, inventory management, and operational efficiency.

AIEdge computingRetail
0 likes · 18 min read
How Low-Cost AI Powers Full-Scale Store Digitalization
DataFunTalk
DataFunTalk
Nov 22, 2020 · Artificial Intelligence

Short Video Analysis in Local Life Scenarios: Techniques and Practices at Meituan

This article presents Meituan's AI-driven short video analysis workflow, covering industry trends, multi‑label video classification, intelligent cover selection, and video generation techniques, while discussing challenges, model building, label expansion, continuous data iteration, and future outlook for video AI in local services.

AIMeituanVideo Generation
0 likes · 16 min read
Short Video Analysis in Local Life Scenarios: Techniques and Practices at Meituan
DeWu Technology
DeWu Technology
Nov 18, 2020 · Artificial Intelligence

AR Fundamentals and Shoe Try‑On Implementation

The presentation explains AR fundamentals, distinguishes it from AI and VR, and details a shoe‑try‑on system that captures 30 fps video, uses AI key‑point detection and pose estimation to overlay 3D shoe models—created via manual, scanning, or photogrammetry methods—rendered with GPU pipelines and PBR, enhanced by green‑screen occlusion and shadow techniques, earning positive audience feedback.

3D modelingARMobile Applications
0 likes · 7 min read
AR Fundamentals and Shoe Try‑On Implementation