Tagged articles

computer vision

687 articles · Page 4 of 7
Python Programming Learning Circle
Python Programming Learning Circle
Mar 8, 2023 · Artificial Intelligence

Using ddddocr SDK for Captcha Recognition in Python

This article introduces the open‑source ddddocr SDK, demonstrates how to install it and use it in Python to automatically solve three common captcha types—slider, click‑based, and alphanumeric—providing code examples and result explanations for each.

OCRcaptchacomputer vision
0 likes · 4 min read
Using ddddocr SDK for Captcha Recognition in Python
Meituan Technology Team
Meituan Technology Team
Feb 23, 2023 · Artificial Intelligence

Food2K: A Large-Scale Food Image Dataset and Progressive Region Enhancement Network

This article reviews the Food2K dataset and the proposed Progressive Region Enhancement Network for large‑scale food image recognition, detailing dataset construction, method design, extensive experiments, ablation studies, visualizations, and future research directions, all validated on the IEEE T‑PAMI 2023 paper.

Fine-Grained ClassificationFood Image RecognitionFood2K
0 likes · 31 min read
Food2K: A Large-Scale Food Image Dataset and Progressive Region Enhancement Network
DaTaobao Tech
DaTaobao Tech
Feb 20, 2023 · Mobile Development

AR Foot Measurement and Hand Try-On Algorithms for Mobile Vision

The article presents a mobile‑vision solution that combines lightweight detection, line detection, segmentation and 3‑D point‑cloud reconstruction to measure foot length within 3 mm error, and a MANO‑based hand‑try‑on system that predicts full mesh vertices for real‑time watch, phone and ring fitting on smartphones.

ARFoot MeasurementHand Try-On
0 likes · 18 min read
AR Foot Measurement and Hand Try-On Algorithms for Mobile Vision
AsiaInfo Technology: New Tech Exploration
AsiaInfo Technology: New Tech Exploration
Feb 20, 2023 · Industry Insights

Why Pre‑trained Large Models Are the New Infrastructure for AI Applications

Pre‑trained large models are emerging as the foundational infrastructure for AI across industries; this article analyzes their technical advantages, application trends in NLP, CV and multimodal domains, presents a telecom customer‑service case study with performance benchmarks, and outlines future deployment challenges and research directions.

Industry ApplicationsNLPPrompt Tuning
0 likes · 23 min read
Why Pre‑trained Large Models Are the New Infrastructure for AI Applications
DataFunTalk
DataFunTalk
Feb 11, 2023 · Artificial Intelligence

Accelerating Computer Vision Pipelines with CV-CUDA: Reducing Complexity and Performance Bottlenecks

This article explains how moving image preprocessing and post‑processing to GPU with the open‑source CV‑CUDA library dramatically reduces system complexity, eliminates CPU‑GPU bottlenecks, and delivers up to thirty‑fold performance gains for computer‑vision workloads across training and inference stages.

CV-CUDAGPU AccelerationPreprocessing
0 likes · 16 min read
Accelerating Computer Vision Pipelines with CV-CUDA: Reducing Complexity and Performance Bottlenecks
DataFunTalk
DataFunTalk
Jan 12, 2023 · Artificial Intelligence

Tencent AI Lab's Advances in High‑Fidelity 3D Face Digitization and Evaluation

This article presents Tencent AI Lab's recent research on efficient 3D face digitization—including single‑photo, multi‑photo, and RGB‑D selfie pipelines—describes a detailed production workflow, introduces a new evaluation benchmark (REALY), and shares insights from a technical Q&A session.

3D face reconstructionAI LabRGBD
0 likes · 11 min read
Tencent AI Lab's Advances in High‑Fidelity 3D Face Digitization and Evaluation
DataFunTalk
DataFunTalk
Jan 8, 2023 · Artificial Intelligence

Adaptive Blend Pyramid Network for Real-Time Local Retouching of Ultra High-Resolution Images

The paper introduces ABPN, an Adaptive Blend Pyramid Network that achieves precise, high‑quality skin retouching and garment wrinkle removal on 4K‑8K photos in real time by combining a context‑aware local retouching layer with a novel adaptive blend pyramid layer, addressing challenges of artifact‑free detail preservation and efficient high‑resolution processing.

adaptive blend pyramidcomputer visiondeep learning
0 likes · 16 min read
Adaptive Blend Pyramid Network for Real-Time Local Retouching of Ultra High-Resolution Images
Kuaishou Audio & Video Technology
Kuaishou Audio & Video Technology
Dec 30, 2022 · Artificial Intelligence

Unlocking Realistic Bokeh: Depth‑Aware Algorithms Behind Holiday Video Effects

This article explains the optical principles of bokeh (scatter blur), describes a depth‑aware variable‑focus algorithm developed by Kuaishou’s audio‑video team, and details practical optimizations such as saliency detection, edge‑preserving weighting, and adaptive spot‑light effects that enable realistic, customizable holiday video filters.

BokehDepth EstimationVideo Effects
0 likes · 11 min read
Unlocking Realistic Bokeh: Depth‑Aware Algorithms Behind Holiday Video Effects
DataFunTalk
DataFunTalk
Dec 27, 2022 · Artificial Intelligence

Efficient Training for Very Large‑Scale Face Recognition and the FFC Framework

This article reviews the challenges of ultra‑large‑scale face recognition, presents existing solutions such as metric learning, PFC and VFC, and details the proposed FFC framework with dual loaders, ID groups, probe and gallery networks, plus experimental results showing its cost‑effective performance.

AIcomputer visiondeep learning
0 likes · 7 min read
Efficient Training for Very Large‑Scale Face Recognition and the FFC Framework
Kuaishou Tech
Kuaishou Tech
Dec 26, 2022 · Artificial Intelligence

ICDAR 2023-DSText Video Text Reading Competition Overview

The ICDAR 2023-DSText competition, launching on February 15, 2023, focuses on dense and small text detection and recognition in video, providing a YouTube‑sourced dataset of 100 videos, two challenge tasks, a detailed timeline, eligibility rules, and a list of international sponsoring institutions.

CompetitionICDARcomputer vision
0 likes · 6 min read
ICDAR 2023-DSText Video Text Reading Competition Overview
Alibaba Cloud Developer
Alibaba Cloud Developer
Dec 19, 2022 · Artificial Intelligence

How AI Transforms Football Video Analysis: Detection, Tracking, and Event Recognition

This article explores how artificial intelligence techniques such as deep learning, object detection, multi‑object tracking, and coordinate projection are applied to football video analysis to automatically detect the ball and players, map their positions onto the field, and recognize key events like shots and goals.

AIcomputer visionobject detection
0 likes · 16 min read
How AI Transforms Football Video Analysis: Detection, Tracking, and Event Recognition
DataFunTalk
DataFunTalk
Dec 17, 2022 · Artificial Intelligence

Multimodal Pre‑training Techniques and Applications – Overview, OPPOVL Dataset, Architecture, and Performance

This article presents a comprehensive overview of multimodal pre‑training, describing its motivation, architecture choices, large‑scale Chinese image‑text dataset construction, training optimizations, performance benchmarks, downstream applications, and a Q&A session that highlights practical deployment considerations.

Large-Scale DataMultimodalcomputer vision
0 likes · 16 min read
Multimodal Pre‑training Techniques and Applications – Overview, OPPOVL Dataset, Architecture, and Performance
Laiye Technology Team
Laiye Technology Team
Dec 16, 2022 · Artificial Intelligence

Efficient Production of Scene-specific OCR Models Using an AI Platform

This article explains how a unified AI platform enables rapid, data‑driven creation, training, deployment, and evaluation of OCR models for visually distinct text regions such as seals, meter readings, license plates, and VIN codes, while minimizing hardware and annotation costs.

AI platformKubeflowOCR
0 likes · 7 min read
Efficient Production of Scene-specific OCR Models Using an AI Platform
DataFunSummit
DataFunSummit
Dec 9, 2022 · Artificial Intelligence

Volcano Engine Virtual Digital Human Technology Overview

This article provides a comprehensive overview of Volcano Engine's virtual digital human platform, detailing its definition, AI‑driven and human‑driven classifications, 2D and 3D technical architectures, multi‑modal perception, interaction capabilities, application scenarios, and future development directions.

2D avatar3D avatarVirtual digital human
0 likes · 15 min read
Volcano Engine Virtual Digital Human Technology Overview
DataFunTalk
DataFunTalk
Dec 5, 2022 · Artificial Intelligence

MogFace: A High‑Performance Face Detector with Dynamic Label Assignment, FP Context Analysis, and Pyramid‑Level Supervision

The article presents MogFace, a state‑of‑the‑art face detection system that combines a dynamic label‑assignment strategy, false‑positive context analysis, and pyramid‑layer ground‑truth supervision to achieve multiple top‑ranked results on the WIDER FACE benchmark, and details its architecture, observations, and experimental validation.

MogFacecomputer visiondynamic label assignment
0 likes · 7 min read
MogFace: A High‑Performance Face Detector with Dynamic Label Assignment, FP Context Analysis, and Pyramid‑Level Supervision
DataFunTalk
DataFunTalk
Nov 17, 2022 · Artificial Intelligence

Enhance the Visual Representation via Discrete Adversarial Training

The Alibaba AAIG team proposes Discrete Adversarial Training (DAT), which leverages VQGAN‑based discretization to generate natural‑looking adversarial samples that improve visual representation robustness and transferability across classification, self‑supervised learning, and object detection tasks without sacrificing accuracy, achieving new state‑of‑the‑art results on multiple benchmarks.

RobustnessVisual Representationadversarial training
0 likes · 12 min read
Enhance the Visual Representation via Discrete Adversarial Training
Tencent Cloud Developer
Tencent Cloud Developer
Nov 11, 2022 · Artificial Intelligence

Tencent Advertising Multimedia AI Technology: Research and Application

Liu Wei outlines Tencent’s Advertising Multimedia AI ecosystem on the Taiji platform, describing a five‑platform matrix—Jue for content understanding, Qiankun for automated video creation, Shenzhen for AI‑driven review, Tianyin for hierarchical fingerprinting, and Hunyuan as a multimodal large model—featuring innovations such as massive multimodal pre‑training, logo retrieval, QA‑style attribute extraction, spatiotemporal video analysis, advanced auto‑judgment, and high‑performance hashing that achieve top cross‑modal retrieval results.

advertising technologycomputer visioncontent understanding
0 likes · 18 min read
Tencent Advertising Multimedia AI Technology: Research and Application
Shopee Tech Team
Shopee Tech Team
Nov 10, 2022 · Artificial Intelligence

ShopeeVideo OCR: Multi-language Text Recognition System for E-commerce Video

ShopeeVideo OCR is a multi‑language text‑recognition system for Southeast Asian e‑commerce videos that unifies detection, Transformer‑based recognition, layout analysis, and large‑scale synthetic data generation to handle Indonesian, Filipino, English, Vietnamese, Thai and Chinese scripts, delivering industry‑leading accuracy and winning thirteen ICDAR first‑place awards.

Multi-language OCROCROptical Character Recognition
0 likes · 15 min read
ShopeeVideo OCR: Multi-language Text Recognition System for E-commerce Video
Rare Earth Juejin Tech Community
Rare Earth Juejin Tech Community
Nov 9, 2022 · Artificial Intelligence

Detailed Explanation of Fully Convolutional Networks (FCN) for Semantic Segmentation

This article provides a comprehensive, beginner‑friendly overview of semantic segmentation, focusing on the pioneering Fully Convolutional Network (FCN) architecture, its variants (FCN‑32s, FCN‑16s, FCN‑8s), underlying concepts, loss computation, and practical tips for working with the VOC dataset.

AlexNetFCNSemantic Segmentation
0 likes · 14 min read
Detailed Explanation of Fully Convolutional Networks (FCN) for Semantic Segmentation
Zhuanzhuan Tech
Zhuanzhuan Tech
Nov 9, 2022 · Artificial Intelligence

Applying OCR to Game Skin Recognition: Filtering Owned Skins and Tolerant Text Matching

This article describes how OCR technology is used in a game marketplace to automatically extract skin parameters from user‑uploaded images, outlines methods for separating owned skin regions from background using color analysis, and presents a tolerant matching solution based on Rabin‑Karp hashing to handle OCR errors.

Game DevelopmentOCRRabin-Karp
0 likes · 10 min read
Applying OCR to Game Skin Recognition: Filtering Owned Skins and Tolerant Text Matching
DataFunSummit
DataFunSummit
Oct 19, 2022 · Artificial Intelligence

Series Six of the Integer Intelligence Autonomous Driving Dataset Collection – Overview and Highlights

This article presents a comprehensive overview of several publicly available autonomous driving datasets, focusing on Series Six of the Integer Intelligence collection, which includes StreetLearn, UTBM RoboCar, Multi‑Vehicle Stereo Event Camera, comma2k19, the Annotated Laser Dataset, Ford, and Oxford RobotCar, detailing their sources, download links, publication years, key features, and research relevance.

autonomous drivingcomputer visiondatasets
0 likes · 10 min read
Series Six of the Integer Intelligence Autonomous Driving Dataset Collection – Overview and Highlights
Baidu Geek Talk
Baidu Geek Talk
Oct 17, 2022 · Artificial Intelligence

OCR Technology: PaddleOCR and Paddle.js Integration

The article explains OCR fundamentals and details how Baidu’s open‑source PaddleOCR suite can be converted and run in browsers via the @paddlejs‑models/ocr SDK, describing model initialization, detection and CRNN‑based recognition pipelines, and presenting benchmark results that show the newer ch_PP‑OCRv2 model achieving higher accuracy and faster inference than the mobile variant.

AIOCRPaddle.js
0 likes · 9 min read
OCR Technology: PaddleOCR and Paddle.js Integration
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Oct 12, 2022 · Artificial Intelligence

Unlock Vision AI: How EasyCV Streamlines Datasets and Model Training

This article introduces EasyCV, an open‑source all‑in‑one visual algorithm platform that abstracts diverse data sources, provides SOTA self‑supervised models, and offers ready‑to‑download datasets for image classification, object detection, segmentation, and pose estimation, complete with configuration examples.

Data pipelinesEasyCVcomputer vision
0 likes · 9 min read
Unlock Vision AI: How EasyCV Streamlines Datasets and Model Training
AntTech
AntTech
Sep 27, 2022 · Artificial Intelligence

Ant Group’s Research Institute Publishes Four NeurIPS 2022 Papers on Advanced Computer Vision and AI

Ant Group’s Ant Technology Research Institute had four papers from its Visual Intelligence Lab accepted at NeurIPS 2022, covering rank diminishing in deep networks, geometry‑aware 3D image synthesis, dynamic discriminators for GANs, and uncertainty‑aware hierarchical refinement for incremental classification, highlighting the institute’s cutting‑edge AI research.

AI researchGANsNeurIPS
0 likes · 8 min read
Ant Group’s Research Institute Publishes Four NeurIPS 2022 Papers on Advanced Computer Vision and AI
Zhengtong Technical Team
Zhengtong Technical Team
Sep 22, 2022 · Artificial Intelligence

How YOLOv5 Powers Real‑Time City Management Video Analysis

This article explains the background, workflow, and technical details of using the YOLOv5 one‑stage object detection algorithm to enable fast, accurate video analytics for urban management, covering data augmentation, backbone design, FPN‑PAN neck, and prediction output processing.

AIYOLOv5city management
0 likes · 8 min read
How YOLOv5 Powers Real‑Time City Management Video Analysis
HomeTech
HomeTech
Sep 20, 2022 · Artificial Intelligence

Deep Learning for Image Classification: Classic Networks, Attention Mechanisms, and Their Application to Fine‑Grained Classification and Automotive Series Recognition

This article reviews the evolution of deep‑learning image‑classification networks, surveys attention mechanisms for fine‑grained tasks, describes the CVPR 2022 FGVC9 competition solution using RegNetY and random attention cropping, and discusses its deployment in automotive series recognition along with future challenges.

Attention MechanismsCVPRFine-Grained Classification
0 likes · 19 min read
Deep Learning for Image Classification: Classic Networks, Attention Mechanisms, and Their Application to Fine‑Grained Classification and Automotive Series Recognition
Programmer DD
Programmer DD
Sep 13, 2022 · Artificial Intelligence

Why AI Porn Detection Still Struggles: Key Challenges Explained

AI-based porn detection uses deep neural networks to classify images, but faces tough hurdles such as visual similarity with benign content, subjective standards for nudity, and vulnerabilities from training‑data dependence, meaning human moderators remain essential for reliable safety.

AI moderationcomputer visioncontent safety
0 likes · 3 min read
Why AI Porn Detection Still Struggles: Key Challenges Explained
DataFunSummit
DataFunSummit
Sep 6, 2022 · Artificial Intelligence

Recent Advances in Self‑Supervised Learning for Text Recognition (OCR)

This article reviews recent progress in applying self‑supervised learning to OCR text recognition, covering mainstream model architectures, key considerations for self‑supervised tasks on text images, and detailed analyses of representative papers such as SeqCLR, SimAN, and DiG, highlighting their designs, experiments, and results.

OCRcomputer visioncontrastive learning
0 likes · 20 min read
Recent Advances in Self‑Supervised Learning for Text Recognition (OCR)
ByteDance Terminal Technology
ByteDance Terminal Technology
Sep 1, 2022 · Artificial Intelligence

Hybrid Computer Vision and Deep Learning for Automated UI Background Color Extraction and Assertion

This article presents a hybrid pipeline combining traditional computer vision techniques and deep learning models to automatically extract and verify text background colors in UI automation screenshots, effectively addressing challenges like limited training data and complex borders to significantly reduce manual inspection costs while achieving high accuracy and robustness in production environments.

UI automationautomated testingcolor extraction
0 likes · 10 min read
Hybrid Computer Vision and Deep Learning for Automated UI Background Color Extraction and Assertion
DevOps
DevOps
Aug 23, 2022 · Artificial Intelligence

Intelligent Automation Testing: Self‑Healing and Machine‑Learning Techniques

This article reviews the evolution of automated testing toward intelligent solutions, explaining self‑healing mechanisms, machine‑learning‑driven object recognition, computer‑vision and OCR approaches, industry tools such as Healenium and Airtest, and future prospects for zero‑code AI‑powered test automation.

AIOCRautomation testing
0 likes · 13 min read
Intelligent Automation Testing: Self‑Healing and Machine‑Learning Techniques
DaTaobao Tech
DaTaobao Tech
Aug 19, 2022 · Artificial Intelligence

SepLUT: Separable Lookup Tables for Real-time Image Enhancement

SepLUT, a new separable lookup‑table framework, splits color enhancement into a 1‑D LUT for independent adjustments and a 3‑D LUT for correlated changes, predicted by a lightweight CNN, enabling quantizable, real‑time ISP performance with state‑of‑the‑art results on the FiveK benchmark.

computer visiondeep learningimage enhancement
0 likes · 12 min read
SepLUT: Separable Lookup Tables for Real-time Image Enhancement
FunTester
FunTester
Aug 18, 2022 · Artificial Intelligence

How AI Can Automate UI Testing: Building Image‑Based Anomaly Detection

This article examines the evolution of mobile UI testing toward AI‑driven approaches, outlines the challenges of large‑scale apps, and details a practical workflow for constructing image‑based anomaly datasets, training a ResNet‑18 model, and iterating on detection performance.

AI testingMobile TestingUI automation
0 likes · 13 min read
How AI Can Automate UI Testing: Building Image‑Based Anomaly Detection
Beike Product & Technology
Beike Product & Technology
Aug 12, 2022 · Artificial Intelligence

Green Area Generation Method Based on Pix2pix Model

This paper proposes a pix2pix‑based method to automatically generate green areas for large‑scale outdoor 3D scene modeling, detailing dataset creation via OpenCV segmentation, model training, region partitioning, and experimental results showing a 93.8% acceptance rate, significantly improving efficiency over manual drawing.

3D modelingGaNcomputer vision
0 likes · 14 min read
Green Area Generation Method Based on Pix2pix Model
AntTech
AntTech
Jul 18, 2022 · Artificial Intelligence

Trusted AI Research at Ant Group: Advances in Computer Vision, Watermark Defense, Robust Machine Learning, and Explainable NLG

Ant Group’s security labs present a series of cutting‑edge AI research achievements—including hierarchical multi‑granular classification for computer vision, watermark‑vaccine defenses, multi‑modal document understanding, robust and explainable machine learning, and logic‑driven data‑to‑text generation—highlighting their commitment to trustworthy and secure AI applications.

AI safetyData2TextRobust Machine Learning
0 likes · 12 min read
Trusted AI Research at Ant Group: Advances in Computer Vision, Watermark Defense, Robust Machine Learning, and Explainable NLG
JD Tech
JD Tech
Jul 18, 2022 · Artificial Intelligence

AI-Powered Visual Defect Detection for Mobile App UI Testing: Methodology, Data Construction, Model Training, and Evaluation

This article presents an end‑to‑end AI‑driven visual testing solution for mobile applications, detailing the business pain points, data set construction, CNN‑based model design, training procedures, performance evaluation with ROC and confusion matrices, and future directions for improving defect detection accuracy.

UI testingcomputer visiondeep learning
0 likes · 14 min read
AI-Powered Visual Defect Detection for Mobile App UI Testing: Methodology, Data Construction, Model Training, and Evaluation
MaGe Linux Operations
MaGe Linux Operations
Jul 14, 2022 · Artificial Intelligence

How to Detect Nude Images with Python and Pillow: A Complete Guide

This article walks through building a Python3 program that uses the Pillow library to identify skin regions in images, applies color‑space heuristics to classify pixels, merges connected skin areas, and decides whether an image is pornographic based on configurable rules, complete with code samples and testing results.

Pythoncomputer visionimage processing
0 likes · 22 min read
How to Detect Nude Images with Python and Pillow: A Complete Guide
58 Tech
58 Tech
Jul 14, 2022 · Artificial Intelligence

Image Quality Assessment Techniques and Their Application in 58.com Recruitment Image Filtering

This article reviews image quality assessment (IQA) methods—including full‑reference, reduced‑reference, and no‑reference approaches—covers typical datasets and evaluation metrics, describes CNN‑based models such as WaDIQaM, DBCNN and hyperIQA, and details a customized IQA solution deployed at 58.com to filter and rank recruitment images, achieving a reduction of bad‑image rate from 9% to 0%.

CNNIQAcomputer vision
0 likes · 17 min read
Image Quality Assessment Techniques and Their Application in 58.com Recruitment Image Filtering
Alimama Tech
Alimama Tech
Jul 13, 2022 · Artificial Intelligence

Fully Automatic Template‑Free Image‑Text Creative Generation System

Alibaba Alimama’s fully automatic, template‑free image‑text creative generation system uses deep‑learning models across material mining, layout synthesis, on‑image copy generation, and visual attribute rendering to produce personalized ad creatives directly from product images and metadata, achieving roughly 19 % CTR lift over prior template‑based methods.

AIAd CreativeAutomation
0 likes · 19 min read
Fully Automatic Template‑Free Image‑Text Creative Generation System
DataFunTalk
DataFunTalk
Jul 12, 2022 · Artificial Intelligence

Applying Computer Vision for Content Safety in Live Streaming: Practices and Future Directions

This presentation details how Huya leverages computer‑vision algorithms to detect and mitigate risky content such as political, pornographic, and violent material in live‑streaming and short‑video platforms, describing system architecture, labeling strategies, algorithmic pipelines, real‑time moderation techniques, and future research directions.

AI safetycomputer visioncontent moderation
0 likes · 11 min read
Applying Computer Vision for Content Safety in Live Streaming: Practices and Future Directions
DaTaobao Tech
DaTaobao Tech
Jul 1, 2022 · Artificial Intelligence

Deep Generative Projection for High‑Fidelity Virtual Try‑On

The paper presents Deep Generative Projection (DGP), a virtual‑try‑on system that learns a realistic dressing distribution from unpaired images with StyleGAN, projects coarse garment‑human alignments into its latent space, refines details, and achieves higher fidelity and robustness than supervised SOTA methods without needing paired data.

Unsupervised LearningVirtual Try-Oncomputer vision
0 likes · 13 min read
Deep Generative Projection for High‑Fidelity Virtual Try‑On
DataFunTalk
DataFunTalk
Jun 30, 2022 · Artificial Intelligence

Self‑Augmented Unpaired Image Dehazing via Density and Depth Decomposition (D4)

The paper introduces D4, a self‑augmented unpaired image dehazing framework that decomposes the transmission map into fog density and scene depth, enabling realistic fog synthesis for data augmentation and achieving superior dehazing performance with fewer parameters and FLOPs on multiple benchmarks.

CVPR2022Depth Estimationcomputer vision
0 likes · 14 min read
Self‑Augmented Unpaired Image Dehazing via Density and Depth Decomposition (D4)
AntTech
AntTech
Jun 24, 2022 · Artificial Intelligence

Hierarchical Residual Network for Multi‑Granularity Classification (HRN) – CVPR 2022 Paper Overview

This article presents a CVPR 2022 paper by Zhejiang University and Ant Group that introduces a label‑relation‑tree‑based Hierarchical Residual Network (HRN) for improving multi‑granularity image classification, detailing its motivation, architecture, composite loss design, extensive experiments on fine‑grained datasets, and practical impact on content‑security applications.

CVPR2022Hierarchical Classificationcomputer vision
0 likes · 12 min read
Hierarchical Residual Network for Multi‑Granularity Classification (HRN) – CVPR 2022 Paper Overview
Meituan Technology Team
Meituan Technology Team
Jun 23, 2022 · Artificial Intelligence

Highlights of Six Meituan Papers Accepted at CVPR 2022

Meituan’s six CVPR 2022 papers advance computer vision by introducing a few‑sample model compression method, a language‑bridged video object segmentation approach, a single‑stage 3D visual grounding technique, a dynamic early‑exit image captioning system, a boosted black‑box adversarial attack, and a semi‑supervised video paragraph grounding framework.

3D groundingCVPR 2022adversarial attacks
0 likes · 15 min read
Highlights of Six Meituan Papers Accepted at CVPR 2022
Xiaohongshu Tech REDtech
Xiaohongshu Tech REDtech
Jun 20, 2022 · Artificial Intelligence

Action Sequence Verification in Videos with CosAlignment Transformer (CAT)

The paper introduces Action Sequence Verification (ASV), a task that determines whether two videos follow the same ordered actions, provides the Chemical Sequence Verification dataset and re‑annotated COIN‑SV and Diving48‑SV sets, and proposes the CosAlignment Transformer (CAT) with intra‑step feature extraction, a Transformer‑based inter‑step encoder, and a sequence‑alignment loss that outperforms prior baselines and serves as a pre‑training model for video retrieval and classification.

Action VerificationMultimodalTransformer
0 likes · 7 min read
Action Sequence Verification in Videos with CosAlignment Transformer (CAT)
Xiaohongshu Tech REDtech
Xiaohongshu Tech REDtech
Jun 13, 2022 · Artificial Intelligence

Neighbor Transformer (NFormer): Robust Person Re-identification via Interactive Multi‑image Modeling

Neighbor Transformer (NFormer) introduces interactive multi‑image modeling for person re‑identification, using Landmark Agent Attention and Reciprocal Neighbor Softmax to efficiently fuse features across images, achieving state‑of‑the‑art accuracy and tighter embedding clusters on multiple benchmark datasets.

computer visiondeep learninglandmark agent attention
0 likes · 8 min read
Neighbor Transformer (NFormer): Robust Person Re-identification via Interactive Multi‑image Modeling
DaTaobao Tech
DaTaobao Tech
Jun 10, 2022 · Artificial Intelligence

NeRF-Editing: Geometry Editing of Neural Radiance Fields

NeRF‑Editing introduces an interactive framework that lets users freely deform the geometry of neural radiance fields by coupling an explicit mesh with implicit NeRF representations, propagating mesh vertex changes through tetrahedral ARAP optimization to bend rays during rendering, enabling realistic edits and animations on synthetic and real‑world scenes, a first reported at CVPR 2022.

3D ReconstructionARAP deformationNeRF
0 likes · 6 min read
NeRF-Editing: Geometry Editing of Neural Radiance Fields
ITPUB
ITPUB
Jun 9, 2022 · Artificial Intelligence

How 58’s Multi‑Label Image Recognition Boosts Semantic Search and Recommendations

This article details the design, data pipeline, model architecture, loss functions, and evaluation metrics of a large‑scale multi‑label image classification system built for 58.com, showing how it improves semantic similarity detection, recommendation, and content moderation across diverse business domains.

Large-Scale Dataasymmetric losscomputer vision
0 likes · 18 min read
How 58’s Multi‑Label Image Recognition Boosts Semantic Search and Recommendations
Python Programming Learning Circle
Python Programming Learning Circle
Jun 9, 2022 · Artificial Intelligence

Python Nude Image Detection Using Pillow: Algorithm, Implementation, and Visualization

This tutorial explains how to build a Python program that detects nude images by analyzing skin-colored regions with Pillow, covering project setup, image preprocessing, pixel classification using RGB/HSV/YCrCb formulas, region merging, decision rules, and command‑line usage with optional visualization.

Nude DetectionPythoncomputer vision
0 likes · 23 min read
Python Nude Image Detection Using Pillow: Algorithm, Implementation, and Visualization
58 Tech
58 Tech
Jun 9, 2022 · Artificial Intelligence

Multi‑Label Image Recognition for 58.com: Algorithm Design, Data Construction, and Model Optimization

This article presents a comprehensive study of multi‑label image recognition applied to 58.com’s business scenarios, covering problem motivation, dataset construction, evaluation metrics, mainstream deep‑learning methods, an asymmetric‑loss‑based optimization pipeline, and practical output schemes for recommendation and retrieval.

asymmetric losscomputer visiondata annotation
0 likes · 17 min read
Multi‑Label Image Recognition for 58.com: Algorithm Design, Data Construction, and Model Optimization
DaTaobao Tech
DaTaobao Tech
Jun 8, 2022 · Artificial Intelligence

Modeling Indirect Illumination for Inverse Rendering

The CVPR‑2022 paper by Alibaba’s Taobao Tech and Zhejiang University introduces a neural‑radiance‑field‑based method that directly models indirect illumination via a signed‑distance‑field geometry and spherical‑Gaussian visibility, avoiding costly path tracing and enabling more accurate recovery of geometry, material and lighting for realistic free‑viewpoint relighting.

BRDFcomputer visionindirect illumination
0 likes · 9 min read
Modeling Indirect Illumination for Inverse Rendering
Youku Technology
Youku Technology
Jun 7, 2022 · Artificial Intelligence

Mobile Real-Time Portrait Segmentation for Youku Bullet Comment Passthrough

To enable real‑time bullet‑comment passthrough on Youku’s mobile app, the team built a million‑scale portrait dataset and designed the AirSegNet series—CPU, GPU, and server variants—using VGG‑style nets, edge‑aware losses, and hybrid CPU‑GPU inference, achieving 0.98 IoU and sub‑15 ms latency on most devices.

Edge computingMNN FrameworkPortrait Segmentation
0 likes · 13 min read
Mobile Real-Time Portrait Segmentation for Youku Bullet Comment Passthrough
NetEase Smart Enterprise Tech+
NetEase Smart Enterprise Tech+
Jun 2, 2022 · Artificial Intelligence

How Knowledge Distillation Shrinks Deep Neural Networks Without Losing Accuracy

Knowledge Distillation, a teacher‑student model compression technique, enables large, high‑performing deep neural networks to transfer their learned representations to smaller models, achieving comparable accuracy with faster inference, reduced resource consumption, and broader applicability in computer‑vision tasks.

AIFitNetcomputer vision
0 likes · 14 min read
How Knowledge Distillation Shrinks Deep Neural Networks Without Losing Accuracy
Java Backend Technology
Java Backend Technology
May 28, 2022 · Artificial Intelligence

5 Mind-Blowing Open-Source Projects That Let You Control Faces, Erase Spiders, and Hack Wi-Fi

This article showcases five cutting‑edge open‑source projects—from a ROS‑based system that lets a gamepad animate facial muscles, to AI‑driven video inpainting, text‑to‑image generation, eye‑gaze computer control, and a comprehensive Wi‑Fi cracking toolkit—each pushing the boundaries of modern tech.

AIcomputer visionnetwork security
0 likes · 6 min read
5 Mind-Blowing Open-Source Projects That Let You Control Faces, Erase Spiders, and Hack Wi-Fi
Youku Technology
Youku Technology
May 18, 2022 · Artificial Intelligence

Subjective and Objective Quality of Experience of Free Viewpoint Videos – Paper Overview

This IEEE TIP paper presents a large‑scale subjective‑objective study of Free Viewpoint Video quality, introducing a cost‑saving two‑stage labeling workflow, a sparse‑frame benchmark model, and publicly releasing the dataset and code, with contributions from Alibaba’s Moku Lab and Jiangxi University researchers.

Free Viewpoint VideoIEEE TIPSparse Sampling
0 likes · 5 min read
Subjective and Objective Quality of Experience of Free Viewpoint Videos – Paper Overview
Code DAO
Code DAO
May 18, 2022 · Artificial Intelligence

A Practical Guide to PyTorch Visualization Tools for Deep Learning

This article walks through the core PyTorch visualization utilities—making image grids, drawing bounding boxes, segmentation masks, and keypoints—explaining why they are needed, how to set up the pipeline, and providing complete code examples for each computer‑vision task.

Bounding BoxesKeypointsPyTorch
0 likes · 18 min read
A Practical Guide to PyTorch Visualization Tools for Deep Learning
DaTaobao Tech
DaTaobao Tech
May 11, 2022 · Artificial Intelligence

AdaInt: Learning Adaptive Intervals for 3D Lookup Tables in Real‑time Image Enhancement

AdaInt introduces a lightweight convolutional network that predicts non‑uniform sampling coordinates and basis 3D LUTs, using a differentiable binary‑search AiLUT‑Transform to enable end‑to‑end training, thereby delivering superior PSNR, negligible extra parameters, and real‑time color enhancement on ultra‑high‑resolution images, outperforming prior state‑of‑the‑art methods.

3D LUTadaptive intervalscomputer vision
0 likes · 11 min read
AdaInt: Learning Adaptive Intervals for 3D Lookup Tables in Real‑time Image Enhancement
Bilibili Tech
Bilibili Tech
May 10, 2022 · Artificial Intelligence

Glance Supervised Video Moment Retrieval via the ViGA Framework

The paper presents a glance‑supervised video moment retrieval approach that records a single annotator‑seen frame, introduces the ViGA contrastive learning framework to leverage this weak temporal cue, and demonstrates on three benchmarks performance rivaling fully supervised methods while keeping annotation cost minimal.

Glance SupervisionMultimodalViGA
0 likes · 8 min read
Glance Supervised Video Moment Retrieval via the ViGA Framework
Code DAO
Code DAO
May 10, 2022 · Artificial Intelligence

How Geometric Deep Learning Enables Spherical CNNs for Rotationally Equivariant Vision

The article explains why traditional planar CNNs fail on spherical data, describes how encoding rotational symmetry through continuous spherical representations and spherical harmonics leads to spherical convolutions that are rotation‑equivariant, and outlines the practical computation using harmonic coefficients.

computer visiongeometric deep learningrotational equivariance
0 likes · 9 min read
How Geometric Deep Learning Enables Spherical CNNs for Rotationally Equivariant Vision
Tencent Cloud Developer
Tencent Cloud Developer
Apr 27, 2022 · Artificial Intelligence

Alignment-Uniformity Representation Learning for Zero-shot Video Classification (AURL)

The AURL framework, presented by Pu Shi, introduces alignment‑uniformity aware representation learning for zero‑shot video classification, achieving up to 28 % top‑1 accuracy gains on UCF101 and HMDB51, and has already boosted business metrics in Tencent’s advertising, search, and video‑channel recommendation systems.

Alignmentcomputer visiondeep learning
0 likes · 19 min read
Alignment-Uniformity Representation Learning for Zero-shot Video Classification (AURL)
Python Programming Learning Circle
Python Programming Learning Circle
Apr 26, 2022 · Artificial Intelligence

Python Script for Adding Face Masks to CelebA Images Using the face_recognition Library

This article demonstrates how to use Python, the face_recognition library, and OpenCV/Pillow to automatically detect facial landmarks in CelebA images, generate and align mask overlays, and save both masked and binary mask versions for computer‑vision research and dataset augmentation.

Pythoncomputer visionface_recognition
0 likes · 11 min read
Python Script for Adding Face Masks to CelebA Images Using the face_recognition Library
Meituan Technology Team
Meituan Technology Team
Apr 14, 2022 · Artificial Intelligence

Short Video Content Understanding and Generation Practices at Meituan

Meituan leverages computer‑vision techniques to tag, analyze, and automatically generate short videos across consumer and merchant scenarios, detailing hierarchical tag design, self‑supervised representation learning, fine‑grained food recognition, intelligent cover creation, and pixel‑level editing to enhance content discovery and presentation.

AI content generationSelf-supervised LearningSemantic Segmentation
0 likes · 20 min read
Short Video Content Understanding and Generation Practices at Meituan
Kuaishou Tech
Kuaishou Tech
Apr 11, 2022 · Artificial Intelligence

Kuaishou's Custom Video Matting Solution: Interactive Object Segmentation for Mobile Creators

Kuaishou's audio‑video technology team presents a self‑developed custom video matting system that combines foreground, interactive, and video object segmentation to let creators extract arbitrary subjects without green screens, featuring adaptive cropping, multi‑stage training, and deployment across Android and iOS devices.

Kuaishoucomputer visiondeep learning
0 likes · 15 min read
Kuaishou's Custom Video Matting Solution: Interactive Object Segmentation for Mobile Creators
Python Programming Learning Circle
Python Programming Learning Circle
Apr 9, 2022 · Artificial Intelligence

Image Resizing with OpenCV and PyTorch

This article explains how to resize images using OpenCV's cv2.resize function and how to scale multi‑dimensional tensors in PyTorch with torch.nn.functional.interpolate, providing detailed parameter descriptions and practical code examples for both single images and batch processing.

PyTorchResizecomputer vision
0 likes · 6 min read
Image Resizing with OpenCV and PyTorch
Meituan Technology Team
Meituan Technology Team
Apr 7, 2022 · Mobile Development

Zero‑Code Scripted Guidance for Mobile Apps Using CV and AI

The ASG system delivers stack‑agnostic, zero‑code in‑app guidance by combining traditional computer‑vision matching with deep‑learning detectors, enabling product teams to author scripts visually, cut development time below half a person‑day, boost task completion from 18 % to 35.7 %, and slash costs over 90 %.

Low-codecomputer visionimage matching
0 likes · 31 min read
Zero‑Code Scripted Guidance for Mobile Apps Using CV and AI
Kuaishou Large Model
Kuaishou Large Model
Apr 6, 2022 · Artificial Intelligence

How Transformers Revolutionize Image Style Transfer: Introducing StyTr²

This article reviews the limitations of traditional CNN‑based image stylization, explains how Transformer architectures overcome these issues with global context and self‑attention, and presents the novel StyTr² method with content‑aware positional encoding that achieves superior, detail‑preserving style transfer results.

Transformercomputer visiondeep learning
0 likes · 8 min read
How Transformers Revolutionize Image Style Transfer: Introducing StyTr²
Tencent Architect
Tencent Architect
Apr 6, 2022 · Artificial Intelligence

Award-Winning AIoT Projects from the 2021 TencentOS Tiny AIoT Innovation Competition

The 2021 TencentOS Tiny AIoT Innovation Competition showcased over 50 original projects, including award‑winning multi‑functional pedestrian detection devices, AI‑enhanced smart wheelchairs, and endangered‑animal recognition systems, each demonstrating low‑power embedded AI, edge computing, and cloud integration for diverse real‑world applications.

AIoTEdge computingEmbedded AI
0 likes · 8 min read
Award-Winning AIoT Projects from the 2021 TencentOS Tiny AIoT Innovation Competition
Kuaishou Tech
Kuaishou Tech
Apr 6, 2022 · Artificial Intelligence

StyTr²: A Transformer‑Based Approach for Image Style Transfer

The paper proposes StyTr², a Transformer‑based image style transfer method that uses content‑aware positional encoding to preserve details and improve feature representation, achieving high‑quality stylization with better content structure and style patterns.

computer visioncontent-aware positional encodingdeep learning
0 likes · 7 min read
StyTr²: A Transformer‑Based Approach for Image Style Transfer
Laiye Technology Team
Laiye Technology Team
Mar 25, 2022 · Artificial Intelligence

Laiye OCR Error‑Correction Model: Architecture, Implementation, and Evaluation

This article describes Laiye's OCR error‑correction system, detailing the background challenges of Chinese character recognition, the analysis of three possible solutions, the chosen post‑processing approach, model architecture, training data, loss design, online inference, and experimental results showing a measurable performance boost.

Chinese textError CorrectionOCR
0 likes · 13 min read
Laiye OCR Error‑Correction Model: Architecture, Implementation, and Evaluation
JD Cloud Developers
JD Cloud Developers
Mar 21, 2022 · Artificial Intelligence

ViTAEv2 Breaks ImageNet Real Record with 91.2% Accuracy – How a 600M‑Parameter Model Redefines Few‑Shot Learning

JD Research Institute and the University of Sydney introduced ViTAEv2, a 600‑million‑parameter deep learning model that achieved a world‑leading 91.2% top‑1 accuracy on ImageNet Real without external data, demonstrating strong few‑shot learning, reducing labeling costs, and promising advances across many computer‑vision tasks.

AI modelImageNetViTAEv2
0 likes · 4 min read
ViTAEv2 Breaks ImageNet Real Record with 91.2% Accuracy – How a 600M‑Parameter Model Redefines Few‑Shot Learning
JD Retail Technology
JD Retail Technology
Mar 7, 2022 · Artificial Intelligence

AI-Driven UI Testing: Data Collection, Model Development, and Deployment for Mobile App Anomaly Detection

This article presents a comprehensive study on applying AI and deep‑learning techniques to mobile UI testing, covering background challenges, feasibility research, abnormal sample construction, model design, training, evaluation, and future directions for intelligent test automation.

AI testinganomaly detectioncomputer vision
0 likes · 13 min read
AI-Driven UI Testing: Data Collection, Model Development, and Deployment for Mobile App Anomaly Detection
Kuaishou Large Model
Kuaishou Large Model
Mar 4, 2022 · Artificial Intelligence

How Adaptive 3D Face Cutout Transforms Kuaishou’s AR Effects

This article explains the adaptive 3D face cutout technology behind Kuaishou's "3D Zoom Face" effect, detailing its problem‑solving approach, implementation workflow, camera‑control optimizations, and how it expands creative possibilities while lowering production costs for both creators and users.

3D renderingAR effectsKuaishou
0 likes · 16 min read
How Adaptive 3D Face Cutout Transforms Kuaishou’s AR Effects
JD Cloud Developers
JD Cloud Developers
Mar 3, 2022 · Artificial Intelligence

How JD Explore’s Silver‑Bullet‑3D Dominated the SAPIEN ManiSkill Challenge

JD Explore Research Institute’s Visual and Multimedia Lab team “Silver‑Bullet‑3D” secured top positions in the 2021 SAPIEN ManiSkill Challenge by excelling in both imitation‑learning and rule‑based tracks, showcasing cutting‑edge computer‑vision and robotic‑arm control technologies that earned them international recognition.

AI competitioncomputer visionimitation learning
0 likes · 5 min read
How JD Explore’s Silver‑Bullet‑3D Dominated the SAPIEN ManiSkill Challenge
Python Crawling & Data Mining
Python Crawling & Data Mining
Feb 22, 2022 · Artificial Intelligence

Create a Dancing Word‑Cloud Video with Python and AI

This tutorial walks through downloading a dance video, extracting frames, using Baidu AI for person segmentation, generating word‑cloud masks, and stitching the results into a dancing word‑cloud video with Python, OpenCV and the WordCloud library.

Baidu AIOpenCVcomputer vision
0 likes · 8 min read
Create a Dancing Word‑Cloud Video with Python and AI
Kuaishou Tech
Kuaishou Tech
Feb 9, 2022 · Mobile Development

Kuaishou Mobile Mixed Reality System: Architecture, Algorithms, and Applications

This article presents Kuaishou's mobile mixed reality (MR) system, detailing its integration of deep learning, SLAM, and scene reconstruction for real‑time spatial computing, the design of a monocular depth‑estimation model, a lightweight 3D rendering engine, and its deployment across iOS and Android devices with various user‑facing effects.

Depth EstimationKuaishouMobile AR
0 likes · 16 min read
Kuaishou Mobile Mixed Reality System: Architecture, Algorithms, and Applications
Baobao Algorithm Notes
Baobao Algorithm Notes
Jan 28, 2022 · Artificial Intelligence

How Masked Autoencoders Revolutionize Vision Pre‑Training: A Deep Dive

This article provides a detailed technical walkthrough of Masked Autoencoders (MAE) for computer vision, covering its BERT‑inspired masking strategy, asymmetric encoder‑decoder design, implementation specifics, experimental findings on mask ratios and decoder depth, and the resulting performance gains over supervised ViT models.

MAEMasked ModelingPyTorch
0 likes · 11 min read
How Masked Autoencoders Revolutionize Vision Pre‑Training: A Deep Dive
Kuaishou Tech
Kuaishou Tech
Jan 27, 2022 · Artificial Intelligence

Kuaishou’s Self‑Developed Green‑Screen Matting Algorithm and Its Deployment in Kuaiying, Live Companion, and Cloud Editing

This article explains the principles, challenges, and implementation details of Kuaishou’s proprietary green‑screen matting algorithm, covering fine‑detail handling, color‑spill reduction, green‑reflection removal, and its real‑time deployment across mobile video‑editing and live‑streaming products.

Kuaishoucomputer visiongreen screen
0 likes · 13 min read
Kuaishou’s Self‑Developed Green‑Screen Matting Algorithm and Its Deployment in Kuaiying, Live Companion, and Cloud Editing
Kuaishou Tech
Kuaishou Tech
Jan 26, 2022 · Artificial Intelligence

Technical Overview of Kuaishou Y‑Tech Body‑Shaping Effects and Underlying Algorithms

This article explains how Kuaishou's Y‑Tech leverages human detection, keypoint localization, and image‑deformation algorithms such as stretching, triangulation and liquify, together with background‑distortion correction, to deliver seven stable, natural body‑shaping effects for short‑video applications.

AIbody shapingcomputer vision
0 likes · 13 min read
Technical Overview of Kuaishou Y‑Tech Body‑Shaping Effects and Underlying Algorithms
Kuaishou Large Model
Kuaishou Large Model
Jan 22, 2022 · Artificial Intelligence

How Kuaishou Achieves Realistic Body Beautification with AI‑Driven Pose Detection and Image Warping

This article explains Kuaishou’s Y‑tech body‑beautification pipeline, detailing how proprietary human pose detection, key‑point localization, and image‑warping techniques such as stretching, triangulation, and liquify are combined to create stable, natural effects like long‑leg, slim‑waist, and swan‑neck, while minimizing background distortion.

AIbody beautificationcomputer vision
0 likes · 15 min read
How Kuaishou Achieves Realistic Body Beautification with AI‑Driven Pose Detection and Image Warping
Baidu Geek Talk
Baidu Geek Talk
Jan 17, 2022 · Artificial Intelligence

Unlocking Video AI: PaddleVideo’s Open‑Source Solutions for Sports, Media, and Safety

This article surveys PaddleVideo, Baidu's open‑source video AI toolkit, detailing its industry‑focused models for sports action recognition, multimodal tagging, intelligent production, interactive segmentation, drone detection, and medical imaging, while providing performance metrics and GitHub resources for each solution.

Open SourcePaddleVideoVideo AI
0 likes · 14 min read
Unlocking Video AI: PaddleVideo’s Open‑Source Solutions for Sports, Media, and Safety
DataFunSummit
DataFunSummit
Jan 5, 2022 · Artificial Intelligence

Improving Financial Micro‑Business Efficiency with OCR: Challenges, Applications, and an Intelligent Platform

This article explores how optical character recognition (OCR) technology can address the financing pain points of micro‑enterprises by automating document verification, enhancing risk assessment, and enabling an end‑to‑end intelligent OCR platform built on deep‑learning models, data pipelines, and deployment automation.

Document AutomationMicro BusinessOCR
0 likes · 15 min read
Improving Financial Micro‑Business Efficiency with OCR: Challenges, Applications, and an Intelligent Platform
Code DAO
Code DAO
Dec 31, 2021 · Artificial Intelligence

Why RegNet Is the Most Flexible Architecture for Computer Vision

RegNet introduces a scalable design space defined by quantized linear functions, enabling flexible trade‑offs between accuracy, efficiency, and mobile deployment, and demonstrates superior performance compared with ResNet, EfficientNet, and other mobile‑optimized networks.

Design SpaceNetwork ArchitectureRegNet
0 likes · 7 min read
Why RegNet Is the Most Flexible Architecture for Computer Vision
Laiye Technology Team
Laiye Technology Team
Dec 31, 2021 · Artificial Intelligence

Overview of Table Recognition Techniques and Practical Implementation

This article reviews the challenges of extracting structured table data from images, compares two‑stage and end‑to‑end OCR approaches, evaluates four state‑of‑the‑art table‑recognition models (SPLERGE, CascadeTabNet, TableMASTER, UnetTable), and presents a practical deployment workflow with performance metrics.

AIOCRTable Recognition
0 likes · 14 min read
Overview of Table Recognition Techniques and Practical Implementation
Code DAO
Code DAO
Dec 29, 2021 · Artificial Intelligence

Understanding Stand-Alone Axial-Attention for Panoptic Segmentation

The paper proposes a stand‑alone axial‑attention mechanism that converts 2‑D attention into 1‑D to lower computational cost while preserving global context, introduces position‑sensitive self‑attention, integrates it into Axial‑ResNet and Axial‑DeepLab, and demonstrates strong results on four large segmentation datasets.

Axial AttentionDeepLabPanoptic Segmentation
0 likes · 7 min read
Understanding Stand-Alone Axial-Attention for Panoptic Segmentation
Laravel Tech Community
Laravel Tech Community
Dec 27, 2021 · Artificial Intelligence

OpenCV 4.5.5 Release Highlights and New Features

OpenCV 4.5.5 introduces audio support in VideoCapture, updates SOVERSION handling, adds OpenVINO 2021.4.2 LTS compatibility, expands ONNX test coverage, upgrades protobuf, optimizes for RISC‑V, and enhances the G‑API module with numerous vectorized kernels, SIMD scheduling, and various bug fixes.

AIG-APIOpenCV
0 likes · 3 min read
OpenCV 4.5.5 Release Highlights and New Features
Code DAO
Code DAO
Dec 22, 2021 · Artificial Intelligence

Understanding SimCLR: A Simple Contrastive Learning Framework for Visual Representations

This article explains SimCLR, the 2020 Google Research framework that advances self‑supervised visual pre‑training by using extensive data augmentations, a ResNet encoder, a projection‑head MLP, and the NT‑Xent loss to learn robust image representations that outperform many prior methods on ImageNet and other benchmarks.

NT-Xent lossResNetSelf-supervised Learning
0 likes · 7 min read
Understanding SimCLR: A Simple Contrastive Learning Framework for Visual Representations
ITPUB
ITPUB
Dec 13, 2021 · Artificial Intelligence

How Data Augmentation Boosts Machine Learning When Data Is Scarce

This article explains how data augmentation can alleviate overfitting by artificially expanding limited training sets, outlines common transformation techniques for images, text, and audio, and discusses the method's benefits, practical applications, and inherent limitations for machine‑learning practitioners.

Data Augmentationcomputer visiondeep learning
0 likes · 6 min read
How Data Augmentation Boosts Machine Learning When Data Is Scarce
Code DAO
Code DAO
Dec 12, 2021 · Artificial Intelligence

Lightning Flash 0.3 Introduces New Tasks, Visualization Tools, Data Pipelines, and Registry API

Lightning Flash 0.3 expands the PyTorch Lightning ecosystem with eight new computer‑vision and NLP tasks, modular API design, integrated model hubs, visualisation callbacks, customizable data‑source hooks, and a central registry for model backbones, all illustrated with concrete code examples.

Data PipelineLightning FlashModel Registry
0 likes · 7 min read
Lightning Flash 0.3 Introduces New Tasks, Visualization Tools, Data Pipelines, and Registry API
Kuaishou Large Model
Kuaishou Large Model
Dec 10, 2021 · Artificial Intelligence

How AI Restores Blurry Faces: Inside Kuaishou’s Y‑Tech High‑Definition Portrait Project

Image clarity impacts daily life, from personal memories to security, and Kuaishou’s Y‑Tech team tackles degradation by constructing paired low‑high quality datasets and a style‑based AI model that leverages facial masks to restore high‑definition portraits, preserving identity while enhancing detail.

AIcomputer visiondeep learning
0 likes · 10 min read
How AI Restores Blurry Faces: Inside Kuaishou’s Y‑Tech High‑Definition Portrait Project
Code DAO
Code DAO
Dec 5, 2021 · Artificial Intelligence

Why DropBlock Outperforms Dropout as an Image Regularizer

This article demonstrates how to implement DropBlock in PyTorch, explains why Dropout fails on image data, details the gamma calculation and mask generation, and shows visual comparisons that illustrate the superiority of contiguous region dropping over random pixel dropout.

DropBlockDropoutPyTorch
0 likes · 11 min read
Why DropBlock Outperforms Dropout as an Image Regularizer