Tagged articles

Deep Learning

1276 articles · Page 4 of 13
DataFunTalk
DataFunTalk
Jun 29, 2024 · Artificial Intelligence

Document Intelligence in the Financial Sector: Technologies, Challenges, and Future Directions

This presentation reviews the technical scope of document intelligence, its specific applications and challenges in finance, recent advances in document analysis, recognition, and understanding, and outlines future research directions for large‑model and multimodal solutions in processing complex financial documents.

Deep LearningDocument AILarge Models
0 likes · 28 min read
Document Intelligence in the Financial Sector: Technologies, Challenges, and Future Directions
JD Tech
JD Tech
Jun 28, 2024 · Artificial Intelligence

An Overview of Large Language Models: History, Fundamentals, Prompt Engineering, Retrieval‑Augmented Generation, Agents, and Multimodal AI

This article provides a comprehensive introduction to large language models, covering their historical development, core architecture, training process, prompt engineering techniques, Retrieval‑Augmented Generation, agent frameworks, multimodal capabilities, safety challenges, and future research directions.

AI AgentsAI safetyDeep Learning
0 likes · 22 min read
An Overview of Large Language Models: History, Fundamentals, Prompt Engineering, Retrieval‑Augmented Generation, Agents, and Multimodal AI
Ops Development & AI Practice
Ops Development & AI Practice
Jun 22, 2024 · Artificial Intelligence

Why Transformers Revolutionized AI: From NLP to Vision and Speech

Transformers, introduced in 2017, have reshaped neural networks by leveraging attention mechanisms to outperform RNNs and CNNs across NLP, computer vision, and speech tasks, offering parallel processing, long‑range dependency capture, and versatile applications such as translation, text generation, image classification, and speech recognition.

Attention MechanismDeep LearningNLP
0 likes · 6 min read
Why Transformers Revolutionized AI: From NLP to Vision and Speech
DataFunTalk
DataFunTalk
Jun 20, 2024 · Artificial Intelligence

User Profiling Algorithms: From Ontology‑Based Methods to Deep Learning and Large Model Integration

This article provides a comprehensive overview of user profiling algorithms, covering the evolution from ontology‑based traditional methods to modern deep‑learning approaches, including structured label prediction, representation learning, active learning, and large‑model integration, while discussing challenges, practical applications, and future research directions.

Active LearningDeep LearningLarge Models
0 likes · 26 min read
User Profiling Algorithms: From Ontology‑Based Methods to Deep Learning and Large Model Integration
AntTech
AntTech
Jun 18, 2024 · Artificial Intelligence

Ant Group’s 24 Papers Featured at CVPR2024: Topics and Abstracts

The IEEE CVPR2024 conference in Seattle accepted 2,719 papers out of 11,532 submissions, and Ant Group contributed 24 papers covering computer vision, deep learning, digital humans, large models, multimodal remote sensing, vision‑language distillation, federated incremental learning, model‑stealing defense, and more, with one highlighted as a highlight.

Ant GroupCVPR2024Deep Learning
0 likes · 17 min read
Ant Group’s 24 Papers Featured at CVPR2024: Topics and Abstracts
Xiaohongshu Tech REDtech
Xiaohongshu Tech REDtech
Jun 17, 2024 · Artificial Intelligence

Xiaohongshu Audio-Video Architecture Team Wins Top Awards in CVPR NTIRE 2024 Challenges

Xiaohongshu’s audio‑video architecture team secured second place in the RAIM challenge and first in the S‑UGC VQA challenge at CVPR NTIRE 2024 by improving generative image restoration with SUPIR, DeSRA and a Fusion model, and enhancing video quality assessment using LIQE, Q‑Align and FAST‑VQA, then deploying these methods for live‑stream denoising, intelligent transcoding and cloud‑based super‑resolution, achieving high PLCC/SROCC scores and up to 33 % bandwidth savings.

AICVPR NTIRE 2024Deep Learning
0 likes · 25 min read
Xiaohongshu Audio-Video Architecture Team Wins Top Awards in CVPR NTIRE 2024 Challenges
Rare Earth Juejin Tech Community
Rare Earth Juejin Tech Community
Jun 16, 2024 · Artificial Intelligence

HRNet Source Code Walkthrough: Keypoint Dataset Construction, Online Data Augmentation, and Training Pipeline

This article provides a detailed, English-language walkthrough of the HRNet source code, covering how the COCO keypoint dataset is built, the online data‑augmentation techniques applied during training, and the end‑to‑end training and inference procedures for human pose estimation.

Data AugmentationDeep LearningHRNet
0 likes · 36 min read
HRNet Source Code Walkthrough: Keypoint Dataset Construction, Online Data Augmentation, and Training Pipeline
Baidu Tech Salon
Baidu Tech Salon
Jun 14, 2024 · Artificial Intelligence

Why Large Models Signal the Dawn of General AI: Insights from Baidu’s CTO

In a keynote at the 2024 Beijing Zhiyuan Conference, Baidu’s CTO Wang Haifeng explained how large‑model universality and comprehensive capabilities are driving artificial general intelligence forward, highlighting scale laws, multimodal advances, agent technologies, and the industrial‑scale production of AI.

AI industrializationAI trendsDeep Learning
0 likes · 7 min read
Why Large Models Signal the Dawn of General AI: Insights from Baidu’s CTO
Rare Earth Juejin Tech Community
Rare Earth Juejin Tech Community
Jun 12, 2024 · Artificial Intelligence

A Simple Introduction to the Transformer Model

This article provides a comprehensive, beginner-friendly explanation of the Transformer architecture, covering its encoder‑decoder structure, self‑attention, multi‑head attention, positional encoding, residual connections, decoding process, final linear and softmax layers, and training considerations, illustrated with numerous diagrams and code snippets.

Deep LearningSelf-AttentionTransformer
0 likes · 24 min read
A Simple Introduction to the Transformer Model
21CTO
21CTO
Jun 2, 2024 · Artificial Intelligence

Geoff Hinton on Scaling Laws, Multimodal AI, and the Future of Intelligence

In a candid interview, Geoff Hinton reflects on his AI journey—from early disappointments in physiology and philosophy to breakthroughs in neural networks, scaling laws, multimodal learning, fast‑weight concepts, and the ethical challenges shaping the future of artificial intelligence.

AI ethicsDeep LearningGeoff Hinton
0 likes · 25 min read
Geoff Hinton on Scaling Laws, Multimodal AI, and the Future of Intelligence
Liangxu Linux
Liangxu Linux
May 26, 2024 · Artificial Intelligence

Can Palette-Based Recoloring Transform Pokémon Images Without Neural Networks?

This article presents a mathematically modeled algorithm that extracts color palettes from any Pokémon image and applies them to another, optimizing the swap via deep‑feature distance and dense color‑transform space, demonstrating superior visual results and subjective evaluations compared to traditional hue‑shift and other recoloring methods.

Deep Learningcolor transfercomputer vision
0 likes · 14 min read
Can Palette-Based Recoloring Transform Pokémon Images Without Neural Networks?
DataFunSummit
DataFunSummit
May 25, 2024 · Artificial Intelligence

Debiased Deep Learning and Double Machine Learning for Multi‑Experiment Causal Inference

This article presents a novel approach that combines debiased deep learning with double machine learning to estimate and infer average treatment effects across multiple simultaneous online experiments, detailing problem definition, a semi‑parametric theoretical framework, and extensive field‑experiment validation on a large video‑platform dataset.

ATE estimationDeep Learningcausal inference
0 likes · 11 min read
Debiased Deep Learning and Double Machine Learning for Multi‑Experiment Causal Inference
Baidu Tech Salon
Baidu Tech Salon
May 24, 2024 · Artificial Intelligence

HelixDock: A Large-Scale Pretrained Full-Atom Diffusion Model for Protein–Small Molecule Docking

HelixDock, a full‑atom diffusion model pretrained on a billion‑scale simulated docking dataset covering ~200,000 protein targets, delivers state‑of‑the‑art docking accuracy—85.6% success on PoseBusters and strong generalization on cross‑docking benchmarks—showing that massive data and model scaling dramatically improve AI‑driven drug discovery, and its code and data are fully open‑source.

AI for drug discoveryDeep LearningHelixDock
0 likes · 6 min read
HelixDock: A Large-Scale Pretrained Full-Atom Diffusion Model for Protein–Small Molecule Docking
Huolala Tech
Huolala Tech
May 23, 2024 · Artificial Intelligence

How to Detect and Remove Moiré Patterns with AI and Diffusion Models

This article explains the nature of moiré patterns in digital imaging, reviews manual mitigation techniques, introduces direct and indirect AI‑based recognition methods—including traditional feature extraction and deep‑learning models such as CNNs and diffusion frameworks—and details practical applications and evaluation metrics used by Huolala.

AIDeep Learningcomputer vision
0 likes · 17 min read
How to Detect and Remove Moiré Patterns with AI and Diffusion Models
Open Source Linux
Open Source Linux
May 22, 2024 · Artificial Intelligence

Why GPUs Are the Powerhouse Behind Modern AI: A Deep Dive

This article explains how GPUs, with their parallel architecture and extensive software ecosystem, have become essential for accelerating AI training and inference, outperforming CPUs and shaping the future of artificial intelligence across various industries.

Artificial IntelligenceDeep LearningGPU
0 likes · 10 min read
Why GPUs Are the Powerhouse Behind Modern AI: A Deep Dive
Architects Research Society
Architects Research Society
May 21, 2024 · Artificial Intelligence

27 Essential AI Papers Recommended by Ilya Sutskever for John Carmack

Ilya Sutskever, former OpenAI chief scientist, shared a curated list of 27 seminal AI research papers—including the Annotated Transformer, Attention Is All You Need, and Deep Residual Learning—with links, claiming mastering them covers roughly 90% of today’s essential artificial‑intelligence knowledge.

AIDeep LearningMachine Learning
0 likes · 7 min read
27 Essential AI Papers Recommended by Ilya Sutskever for John Carmack
Baidu Tech Salon
Baidu Tech Salon
May 20, 2024 · Artificial Intelligence

HelixFold-Multimer: High‑Performance Antigen‑Antibody and Peptide‑Protein Complex Structure Prediction

HelixFold‑Multimer, a new Baidu PaddleHelix model, outperforms AlphaFold 3 on antigen‑antibody and peptide‑protein complex predictions, achieving mean DockQ scores of 0.41 and 0.38 respectively and success rates up to 77 % when epitope data are used, and is already deployed in large‑molecule drug pipelines.

BioinformaticsDeep LearningHelixFold-Multimer
0 likes · 7 min read
HelixFold-Multimer: High‑Performance Antigen‑Antibody and Peptide‑Protein Complex Structure Prediction
JD Tech
JD Tech
May 17, 2024 · Artificial Intelligence

Optimizing JD Advertising Retrieval Platform: Balancing Compute, Data Scale, and Iterative Efficiency

The article details how JD's advertising retrieval platform tackles the core challenge of balancing limited compute resources with massive data by optimizing compute allocation, improving model scoring efficiency, and enhancing iteration speed through distributed execution graphs, adaptive algorithms, and platform‑level infrastructure improvements.

ANNAdvertisingDeep Learning
0 likes · 24 min read
Optimizing JD Advertising Retrieval Platform: Balancing Compute, Data Scale, and Iterative Efficiency
Architects' Tech Alliance
Architects' Tech Alliance
May 14, 2024 · Artificial Intelligence

Why GPUs Are Essential for Modern Artificial Intelligence and How They Compare with CPUs, ASICs, and FPGAs

This article explains the pivotal role of GPUs in today’s generative AI era, describes their architecture and applications, compares them with CPUs, ASICs, and FPGAs, and offers guidance on selecting the right processor for AI workloads while also noting related reference resources.

Artificial IntelligenceDeep LearningGPU
0 likes · 12 min read
Why GPUs Are Essential for Modern Artificial Intelligence and How They Compare with CPUs, ASICs, and FPGAs
Architect's Guide
Architect's Guide
May 13, 2024 · Artificial Intelligence

Understanding the Core Principles of Transformer Architecture

This article explains how Transformer models work by detailing the encoder‑decoder structure, self‑attention, multi‑head attention, positional encoding, and feed‑forward networks, and shows their applications in machine translation, recommendation systems, and large language models.

AIAttention MechanismDeep Learning
0 likes · 11 min read
Understanding the Core Principles of Transformer Architecture
JD Cloud Developers
JD Cloud Developers
Apr 30, 2024 · Artificial Intelligence

Build a Handwritten Digit Recognizer in Java with TensorFlow

This article walks through the complete process of creating, training, evaluating, saving, and loading a MNIST handwritten digit recognition model using TensorFlow in Java, comparing it with the equivalent Python implementation and covering required knowledge, environment setup, and code details.

Deep LearningJavaMNIST
0 likes · 34 min read
Build a Handwritten Digit Recognizer in Java with TensorFlow
Rare Earth Juejin Tech Community
Rare Earth Juejin Tech Community
Apr 24, 2024 · Artificial Intelligence

Training MNIST with Burn on wgpu: From PyTorch to Rust Backend

This tutorial demonstrates how to train a MNIST digit‑recognition model using the Rust‑based Burn framework on top of the cross‑platform wgpu API, covering model export from PyTorch to ONNX, code generation, data loading, training loops, and performance comparison across CPU, GPU, and other backends.

BurnDeep LearningGPU
0 likes · 13 min read
Training MNIST with Burn on wgpu: From PyTorch to Rust Backend
DaTaobao Tech
DaTaobao Tech
Apr 22, 2024 · Artificial Intelligence

Neural Networks and Deep Learning: Principles and MNIST Example

The article reviews recent generative‑AI breakthroughs such as GPT‑5 and AI software engineers, explains that AI systems are deterministic rather than black boxes, and then teaches neural‑network fundamentals—including activation functions, back‑propagation, and a hands‑on MNIST digit‑recognition example with discussion of overfitting and regularization.

Deep LearningMNISTactivation functions
0 likes · 17 min read
Neural Networks and Deep Learning: Principles and MNIST Example
Top Architect
Top Architect
Apr 18, 2024 · Artificial Intelligence

Understanding Transformers: Architecture, Attention Mechanism, Training and Inference

This article provides a comprehensive overview of Transformer models, covering their attention-based architecture, encoder-decoder structure, training procedures including teacher forcing, inference workflow, advantages over RNNs, and various applications in natural language processing such as translation, summarization, and classification.

Attention MechanismDeep LearningNLP
0 likes · 11 min read
Understanding Transformers: Architecture, Attention Mechanism, Training and Inference
NewBeeNLP
NewBeeNLP
Apr 16, 2024 · Artificial Intelligence

Demystifying the Transformer: Step‑by‑Step PaddlePaddle Implementation

This article provides a comprehensive, code‑rich walkthrough of the Transformer architecture using PaddlePaddle, covering the encoder and decoder components, residual connections, layer normalization, feed‑forward networks, scaled dot‑product and multi‑head attention, and shows how to assemble the full model with training and inference functions.

Attention MechanismDeep LearningEncoder
0 likes · 17 min read
Demystifying the Transformer: Step‑by‑Step PaddlePaddle Implementation
DataFunSummit
DataFunSummit
Apr 15, 2024 · Artificial Intelligence

Deep Learning Practices for Internet Real‑Estate Recommendation at 58.com

This article details the end‑to‑end deep‑learning pipeline used by 58.com for real‑estate recommendation, covering business background, a six‑layer architecture, vector‑based recall, various embedding and ranking models, multi‑task and multi‑scenario optimization techniques, and future directions for large‑model integration.

Deep LearningFAISSmulti-task learning
0 likes · 19 min read
Deep Learning Practices for Internet Real‑Estate Recommendation at 58.com
AI Algorithm Path
AI Algorithm Path
Apr 5, 2024 · Artificial Intelligence

Master CNN, RNN, GAN, and Transformer Architectures in One Guide

This article provides a friendly, step‑by‑step overview of five core deep‑learning architectures—CNN, RNN, GAN, Transformers, and encoder‑decoder—explaining their structures, key components, and typical use cases in image and natural‑language processing.

CNNDeep LearningEncoder-Decoder
0 likes · 12 min read
Master CNN, RNN, GAN, and Transformer Architectures in One Guide
DataFunTalk
DataFunTalk
Apr 2, 2024 · Artificial Intelligence

User Portrait Algorithms: From Ontology‑Based Methods to Deep Learning and Future Directions

This article provides a comprehensive overview of user portrait algorithms, covering their historical development, ontology‑based traditional approaches, deep‑learning enhancements, representation‑learning techniques such as lookalike, active‑learning driven iteration, and the integration of large‑model world knowledge, while also discussing current challenges and future research directions.

Active LearningDeep LearningLarge Language Models
0 likes · 26 min read
User Portrait Algorithms: From Ontology‑Based Methods to Deep Learning and Future Directions
Architect
Architect
Mar 28, 2024 · Artificial Intelligence

Understanding OpenAI's Sora Video Generation Model: Architecture, Workflow, and Core Technologies

This article explains OpenAI's Sora video generation model, detailing its latent diffusion foundation, video compression network, spacetime patch representation, Diffusion Transformer processing, and decoding pipeline, while also reviewing related Stable Diffusion and Transformer concepts that enable high‑quality text‑to‑video synthesis.

AIDeep LearningSora
0 likes · 17 min read
Understanding OpenAI's Sora Video Generation Model: Architecture, Workflow, and Core Technologies
Test Development Learning Exchange
Test Development Learning Exchange
Mar 27, 2024 · Artificial Intelligence

Introduction to PyTorch and Example CNN Training on CIFAR-10

This article introduces PyTorch as a leading open‑source deep‑learning framework, outlines its key components such as dynamic computation graphs, tensors, autograd, modules, optimizers, data loading, distributed training and TorchScript, and provides a complete Python example that defines a simple CNN and trains it on the CIFAR‑10 dataset.

CNNDeep LearningPyTorch
0 likes · 8 min read
Introduction to PyTorch and Example CNN Training on CIFAR-10
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Mar 15, 2024 · Artificial Intelligence

Why Arithmetic Feature Interaction Is Key to Deep Tabular Learning

Researchers from Alibaba Cloud AI and Zhejiang University present AMFormer, a Transformer‑based model that incorporates arithmetic feature interaction, demonstrating superior fine‑grained modeling, sample efficiency, and generalization on synthetic and real‑world tabular datasets, establishing a new state‑of‑the‑art in deep tabular learning.

AMFormerDeep LearningMachine Learning
0 likes · 12 min read
Why Arithmetic Feature Interaction Is Key to Deep Tabular Learning
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Mar 12, 2024 · Artificial Intelligence

AAAI‑2024 Highlights: Alibaba Cloud’s Deep Tabular Learning & Multi‑Modal Fusion

Alibaba Cloud’s AI platform PAI showcased four cutting‑edge papers at AAAI‑2024—introducing AMFormer for deep tabular learning via arithmetic feature interaction, MuLTI for efficient video‑language understanding, M2SD for few‑shot class‑incremental learning, and M2Doc for multi‑modal document layout analysis—demonstrating the platform’s growing impact on artificial‑intelligence research.

Deep LearningMultimodal AIdocument analysis
0 likes · 9 min read
AAAI‑2024 Highlights: Alibaba Cloud’s Deep Tabular Learning & Multi‑Modal Fusion
Sohu Tech Products
Sohu Tech Products
Mar 6, 2024 · Artificial Intelligence

Mastering Regression: A Comprehensive Guide to Linear and Non‑Linear Models

This article provides an in‑depth overview of regression prediction, covering linear models like OLS, Lasso, Ridge, and Bayesian approaches, as well as non‑linear techniques such as tree ensembles, SVR, KNN, neural networks, and advanced deep learning frameworks for tabular data.

Deep LearningMachine Learninggradient boosting
0 likes · 13 min read
Mastering Regression: A Comprehensive Guide to Linear and Non‑Linear Models
NewBeeNLP
NewBeeNLP
Mar 4, 2024 · Artificial Intelligence

A Curated Tour of Mamba Papers: 25 Cutting‑Edge State‑Space Model Innovations

This article presents a GitHub‑hosted collection of 25 recent research papers on Mamba and its variants, summarizing each work’s core contributions across sequence modeling, vision, medical imaging, graph analysis, and multimodal tasks, and highlighting their performance gains over prior methods.

Deep LearningMambaSequence Modeling
0 likes · 13 min read
A Curated Tour of Mamba Papers: 25 Cutting‑Edge State‑Space Model Innovations
Bilibili Tech
Bilibili Tech
Mar 1, 2024 · Artificial Intelligence

Bilibili's Self-Developed Video Super-Resolution Algorithm: Background, Optimization Directions, and Implementation Details

Bilibili’s self‑supervised video super‑resolution system upgrades low‑resolution streams to 4K by using three parallel degradation‑branch networks—texture‑enhancing, line‑recovering, and noise‑removing—tailored to anime, game, and real‑world content, delivering sharper edges, finer textures, and measurable quality gains across its online playback pipeline.

AIBilibiliDeep Learning
0 likes · 16 min read
Bilibili's Self-Developed Video Super-Resolution Algorithm: Background, Optimization Directions, and Implementation Details
DataFunTalk
DataFunTalk
Feb 24, 2024 · Artificial Intelligence

Causal Learning Paradigms: From Prior Causal Structure to Causal Discovery

This article introduces causal learning, explains its distinction from traditional correlation‑based machine learning, outlines its three main parts, discusses the two primary paradigms—learning with known causal graphs and learning via causal discovery—and highlights their advantages, challenges, and recent research directions.

Deep LearningDomain AdaptationMachine Learning
0 likes · 11 min read
Causal Learning Paradigms: From Prior Causal Structure to Causal Discovery
Ximalaya Technology Team
Ximalaya Technology Team
Feb 20, 2024 · Artificial Intelligence

Optimization of Deep Learning-Based CTR Models in Advertising

This report presents recent advances in optimizing deep learning click‑through‑rate models for advertising, including improved embedding mechanisms, novel feature‑interaction and architecture designs such as attention‑based behavior sequencing, multi‑tower and Mixture‑of‑Experts networks, dynamic ID handling, hourly updates, incremental training, and outlines future multi‑modal and embedding‑importance research.

Attention MechanismsCTR modelDeep Learning
0 likes · 13 min read
Optimization of Deep Learning-Based CTR Models in Advertising
21CTO
21CTO
Feb 18, 2024 · Artificial Intelligence

How OpenAI’s Sora Turns Text into Realistic 60‑Second Videos

OpenAI’s newly unveiled Sora system can generate 60‑second, high‑quality videos from plain text prompts, leveraging a data‑driven physical engine trained on synthetic data from Unreal Engine 5, with contributions from researchers like Tim Brooks and Bill Peebles, marking a major AI video‑generation breakthrough.

Deep LearningOpenAIgenerative AI
0 likes · 6 min read
How OpenAI’s Sora Turns Text into Realistic 60‑Second Videos
DataFunSummit
DataFunSummit
Feb 9, 2024 · Artificial Intelligence

STAN: A User‑Lifecycle‑Based Multi‑Task Recommendation Model for Shopee

The article introduces STAN, a multi‑task recommendation framework that leverages user lifecycle segmentation to jointly optimize CTR, stay‑time, and CVR, detailing the business context, key challenges, solution architecture, offline and online evaluations, and future research directions.

CTRCVRDeep Learning
0 likes · 8 min read
STAN: A User‑Lifecycle‑Based Multi‑Task Recommendation Model for Shopee
Baidu Geek Talk
Baidu Geek Talk
Feb 5, 2024 · Artificial Intelligence

Why Static Graphs Outperform Dynamic Graphs in AutoDiff: A Deep Dive

This article explains the fundamental differences between static and dynamic computation graphs, compares their memory and performance characteristics, shows how automatic differentiation works in each paradigm, and provides a step‑by‑step implementation of a toy static‑graph AutoDiff engine with Python code examples.

AutoDiffDeep LearningDynamic Graph
0 likes · 18 min read
Why Static Graphs Outperform Dynamic Graphs in AutoDiff: A Deep Dive
360 Smart Cloud
360 Smart Cloud
Jan 26, 2024 · Artificial Intelligence

Parallel Strategies for Distributed Deep Learning Training

This article reviews distributed training techniques for large deep‑learning models, covering data parallelism, model parallelism (including pipeline and tensor parallelism), gradient bucketing and accumulation, 3D parallelism, and practical implementations such as Megatron‑LM and 360AI platform optimizations.

AIDeep LearningMegatron-LM
0 likes · 22 min read
Parallel Strategies for Distributed Deep Learning Training
DataFunSummit
DataFunSummit
Jan 1, 2024 · Artificial Intelligence

Advances in Image and Video Enhancement, Quality Assessment, and Multimodal AI Techniques

This article reviews the latest research from Alibaba DAMO Academy on real-world image quality problems, covering spatial, temporal, and color enhancement methods, advanced quality assessment metrics, multimodal diffusion models, and future directions toward large‑model integration and lightweight deployment.

Deep LearningMOS regressionMultimodal AI
0 likes · 24 min read
Advances in Image and Video Enhancement, Quality Assessment, and Multimodal AI Techniques
Sohu Tech Products
Sohu Tech Products
Dec 27, 2023 · Artificial Intelligence

Analysis of LLaMA Model Architecture in the Transformers Library

This article walks through the core LLaMA implementation in HuggingFace’s Transformers library, detailing the inheritance hierarchy, configuration defaults, model initialization, embedding and stacked decoder layers, the RMSNorm‑based attention and MLP modules, and the forward pass that produces normalized hidden states.

Artificial IntelligenceDeep LearningPyTorch
0 likes · 14 min read
Analysis of LLaMA Model Architecture in the Transformers Library
Huolala Tech
Huolala Tech
Dec 26, 2023 · Artificial Intelligence

How AI Powers Scalable Multilingual and Timezone Testing for Global Apps

This article explains how a deep‑learning‑driven AI platform tackles the complex challenges of multilingual and multi‑timezone testing for a rapidly expanding international app, detailing the architecture, data pipelines, model training, and the resulting efficiency, accuracy, and coverage gains.

AIDeep LearningMultilingual Testing
0 likes · 14 min read
How AI Powers Scalable Multilingual and Timezone Testing for Global Apps
Bilibili Tech
Bilibili Tech
Dec 15, 2023 · Artificial Intelligence

Bilibili's AI-Powered Video Frame Interpolation: Techniques, Challenges, and Deployment

Bilibili’s AI‑driven frame‑interpolation pipeline upgrades low‑frame-rate videos to smooth high‑frame-rate 1080p playback by optimizing optical‑flow models for large motion, texture and text artifacts, pruning for speed, and deploying via the BVT SDK across on‑demand and live streams.

AIDeep LearningVideo Frame Interpolation
0 likes · 14 min read
Bilibili's AI-Powered Video Frame Interpolation: Techniques, Challenges, and Deployment
DataFunTalk
DataFunTalk
Dec 10, 2023 · Artificial Intelligence

PyTorch Model Training Performance Tuning Guide

This guide provides comprehensive techniques for optimizing PyTorch training performance and efficiency, covering all model types such as CNNs, RNNs, GANs, and transformers, and applicable across domains like computer vision and natural language processing, targeting AI/ML platform engineers, data engineers, backend developers, MLOps, SREs, architects, and machine learning engineers.

AIDeep LearningPerformance Tuning
0 likes · 2 min read
PyTorch Model Training Performance Tuning Guide
DataFunSummit
DataFunSummit
Dec 9, 2023 · Artificial Intelligence

Causal Learning Paradigms: From Prior Causal Structure to Causal Discovery

This article reviews the growing interest in causal learning within machine learning, explaining what causal learning is, its advantages over purely correlational methods, and detailing two main paradigms—learning with known causal structures and learning via causal discovery—along with examples, challenges, and future directions.

Deep LearningDomain AdaptationMachine Learning
0 likes · 12 min read
Causal Learning Paradigms: From Prior Causal Structure to Causal Discovery
Airbnb Technology Team
Airbnb Technology Team
Dec 8, 2023 · Artificial Intelligence

Leveraging Image Aesthetics and Photo Sorting Algorithms to Enhance Airbnb Listings

Airbnb’s new computer‑vision pipeline trains a deep‑learning aesthetic model with an EMD loss to rank photos, automatically sorts new‑listing images by design and room type, and scales real‑time similarity search via HNSW‑based ANN on AWS OpenSearch, boosting click‑through, bookings, and enabling unsupervised visual recommendations.

AirbnbDeep LearningEmbedding Search
0 likes · 9 min read
Leveraging Image Aesthetics and Photo Sorting Algorithms to Enhance Airbnb Listings
Rare Earth Juejin Tech Community
Rare Earth Juejin Tech Community
Dec 8, 2023 · Artificial Intelligence

Simplifying Transformer Blocks: Removing Residual Connections, LayerNorm, and Other Components without Losing Performance

A recent ETH Zurich paper shows that standard Transformer blocks can be drastically simplified by removing residual connections, LayerNorm, projection and value parameters, and even MLP sub‑block components, achieving up to 16% fewer parameters and comparable training speed and downstream performance on both GPT‑style decoders and BERT models.

AIDeep LearningLLM
0 likes · 11 min read
Simplifying Transformer Blocks: Removing Residual Connections, LayerNorm, and Other Components without Losing Performance
IT Services Circle
IT Services Circle
Dec 6, 2023 · Artificial Intelligence

AI Image Outpainting: Unexpected Transformations and How It Works

The article showcases a series of humorous and surprising AI‑generated image expansions from Douyin, explains the underlying outpainting technology, and discusses why such tools are both entertaining and useful despite occasional odd results.

AIDeep LearningGenerative Fill
0 likes · 6 min read
AI Image Outpainting: Unexpected Transformations and How It Works
Huolala Tech
Huolala Tech
Nov 28, 2023 · Mobile Development

How HuoLala Built a Low‑Cost, High‑Reliability Mobile UI Automation Platform

This article details HuoLala's journey from a weekly release cycle to a cloud‑based record‑and‑replay mobile UI automation platform, covering background challenges, industry analysis, technical design—including deep‑learning based control detection, SIFT image matching, script generation, playback handling, and platform features—while demonstrating significant testing efficiency gains and future AI‑driven enhancements.

CI/CDDeep LearningRecord & Replay
0 likes · 21 min read
How HuoLala Built a Low‑Cost, High‑Reliability Mobile UI Automation Platform
dbaplus Community
dbaplus Community
Nov 27, 2023 · Artificial Intelligence

Build an Image‑Search Engine with Elasticsearch 8.x and CLIP

This guide explains how to implement reverse image search by extracting visual features with a multilingual CLIP model, storing the vectors in Elasticsearch 8.x, and using its k‑NN plugin to retrieve similar images, covering architecture, tools, code snippets, and results.

CLIPDeep Learningimage search
0 likes · 9 min read
Build an Image‑Search Engine with Elasticsearch 8.x and CLIP
JD Retail Technology
JD Retail Technology
Nov 23, 2023 · Artificial Intelligence

Recent Advances in Advertising Recommendation Algorithms and Their Applications

This article reviews recent progress in advertising recommendation technologies, covering deep learning‑based ranking, sequence modeling, self‑supervised learning, online and reinforcement learning, multimodal recommendation, and fairness, and details four key breakthroughs—data‑driven incremental learning, dynamic group parameter modeling, bilateral interactive graph convolution, and a relation‑aware diffusion model for poster layout generation, along with experimental results and future challenges.

Deep Learningadvertising recommendationdiffusion models
0 likes · 25 min read
Recent Advances in Advertising Recommendation Algorithms and Their Applications
Rare Earth Juejin Tech Community
Rare Earth Juejin Tech Community
Nov 15, 2023 · Artificial Intelligence

Understanding the Transformer Architecture: Encoder, Decoder, and Attention Mechanisms

This article explains the Transformer model, comparing it with RNNs, detailing its encoder‑decoder structure, multi‑head and scaled dot‑product attention, embedding layers, feed‑forward networks, and the final linear‑softmax output, supplemented with diagrams and code examples.

Artificial IntelligenceDeep LearningEncoder-Decoder
0 likes · 10 min read
Understanding the Transformer Architecture: Encoder, Decoder, and Attention Mechanisms
Network Intelligence Research Center (NIRC)
Network Intelligence Research Center (NIRC)
Nov 9, 2023 · Artificial Intelligence

How Wav2Lip Achieves Accurate Speech‑Driven Lip Sync with Expert Discriminators

The article analyzes the limitations of traditional speech‑driven lip‑sync methods and explains how Wav2Lip introduces a pretrained multi‑frame expert sync discriminator, a two‑stage GAN training pipeline, and a specialized generator architecture to produce high‑quality, audio‑aligned facial videos.

Deep LearningGaNWav2Lip
0 likes · 7 min read
How Wav2Lip Achieves Accurate Speech‑Driven Lip Sync with Expert Discriminators
NetEase Media Technology Team
NetEase Media Technology Team
Nov 6, 2023 · Artificial Intelligence

Overview of Sequential Recommendation Models

The article surveys sequential recommendation models from early non-deep approaches like FPMC, through RNN-based GRU4Rec and CNN-based Caser, to Transformer-based methods such as SASRec, BERT4Rec, TiSASRec, and recent contrastive-learning techniques, recommending SASRec or its variants for production use.

Deep LearningTransformercontrastive learning
0 likes · 17 min read
Overview of Sequential Recommendation Models
Python Programming Learning Circle
Python Programming Learning Circle
Oct 26, 2023 · Artificial Intelligence

Animal Recognition Techniques Using Deep Learning and Image Processing

This article reviews animal recognition technology, covering its background, basic principles, image‑processing, feature extraction, machine‑learning and deep‑learning methods, dataset construction, preprocessing, and feature‑selection techniques, and provides Python code examples for implementing CNNs and traditional classifiers.

Deep LearningMachine Learninganimal recognition
0 likes · 18 min read
Animal Recognition Techniques Using Deep Learning and Image Processing
Model Perspective
Model Perspective
Oct 18, 2023 · Fundamentals

Unlock the Power of Convolution: From Signal Smoothing to Deep Learning

This article explains the mathematical definition of convolution, walks through discrete and continuous examples, demonstrates its use in signal smoothing with moving averages, and surveys its wide-ranging applications in signal processing, communications, computer vision, seismology, medical imaging, and statistics.

ConvolutionDeep LearningMathematics
0 likes · 7 min read
Unlock the Power of Convolution: From Signal Smoothing to Deep Learning
DataFunSummit
DataFunSummit
Oct 17, 2023 · Artificial Intelligence

DataFunSummit2023: Deep Learning‑Driven Multi‑Experiment Causal Inference and Distributed Causal Tools

The DataFunSummit2023 online conference brings together experts from Tencent and Kuaishou to present cutting‑edge research on causal inference for large‑scale A/B testing, including deep‑learning‑based multi‑experiment effect estimation, a distributed causal inference framework (Fast‑Causal‑Inference), and strategies for evaluating long‑term policy impacts.

A/B testingDeep LearningMachine Learning
0 likes · 7 min read
DataFunSummit2023: Deep Learning‑Driven Multi‑Experiment Causal Inference and Distributed Causal Tools
Kuaishou Tech
Kuaishou Tech
Oct 17, 2023 · Artificial Intelligence

QIN: A Query‑Dominated User Interest Network for Personalized Search

The paper introduces QIN, a query‑driven user interest network that combines a Relevance Search Unit and a Fused Attention Unit to effectively leverage full‑history user behavior for personalized search, demonstrating significant performance gains in offline benchmarks and online A/B tests.

Deep Learningfused attentionpersonalized search
0 likes · 9 min read
QIN: A Query‑Dominated User Interest Network for Personalized Search
Meituan Technology Team
Meituan Technology Team
Oct 11, 2023 · Artificial Intelligence

Meituan Vision AI Research Highlights and Open‑Source Releases

This article compiles Meituan's cutting‑edge computer‑vision research and engineering achievements—including CVPR award‑winning segmentation, YOLOv6 releases, GPU inference optimizations, the Food2K dataset, and numerous paper digests—to provide practical insights for visual AI practitioners.

CVPRDeep LearningFood2K
0 likes · 11 min read
Meituan Vision AI Research Highlights and Open‑Source Releases
Architect
Architect
Oct 4, 2023 · Artificial Intelligence

How AI-Driven Digital Watermarks Achieve Robust, Invisible Protection for Video

This article examines the challenges of video copyright protection, critiques traditional visible and invisible watermark methods, and presents a deep‑learning based AI digital watermark solution that balances invisibility and robustness, detailing its network architecture, degradation layer, loss functions, block encoding, anchor calibration, and large‑scale experimental results.

AI video protectionDeep LearningRobustness
0 likes · 22 min read
How AI-Driven Digital Watermarks Achieve Robust, Invisible Protection for Video
DataFunSummit
DataFunSummit
Oct 3, 2023 · Artificial Intelligence

Time Series Forecasting for NIO Power Swap Stations: Business Background, Challenges, Algorithm Practice, and Future Outlook

This article presents a comprehensive case study of NIO's Power swap‑station ecosystem, detailing the business context, key forecasting challenges, the evolution from classical statistical models to deep‑learning architectures with specialized embeddings, and the practical outcomes and future plans for improving prediction accuracy.

Deep LearningElectric VehicleMachine Learning
0 likes · 16 min read
Time Series Forecasting for NIO Power Swap Stations: Business Background, Challenges, Algorithm Practice, and Future Outlook
DataFunSummit
DataFunSummit
Sep 29, 2023 · Artificial Intelligence

Social4Rec: Enhancing Video Recommendation with Social Interest Networks

This article introduces Social4Rec, a video recommendation algorithm that tackles user cold‑start problems by extracting and integrating social interest information through coarse‑ and fine‑grained interest extractors, attention‑based fusion, and extensive offline and online experiments demonstrating significant CTR improvements.

Deep Learningattentioncold-start
0 likes · 14 min read
Social4Rec: Enhancing Video Recommendation with Social Interest Networks
Bilibili Tech
Bilibili Tech
Sep 29, 2023 · Artificial Intelligence

BILIVQA: Bilibili's No-Reference Video Quality Assessment System

BILIVQA is Bilibili’s deep‑learning, no‑reference video quality assessment system that trains on a proprietary 5,000‑video UGC dataset, extracts spatial and temporal features via MobileNet‑V2 and X3D, uses mixed‑dataset regression for strong generalization, and deploys a GPU‑optimized TensorRT pipeline with percentile‑based scoring for reliable quality monitoring and downstream applications.

BILIVQADeep LearningModel Engineering
0 likes · 27 min read
BILIVQA: Bilibili's No-Reference Video Quality Assessment System
Zhuanzhuan Tech
Zhuanzhuan Tech
Sep 28, 2023 · Artificial Intelligence

Evolution of Language Models and an Overview of the GPT Series

This article surveys the development of natural language processing from early rule‑based systems through statistical n‑gram models, neural language models, RNNs, LSTMs, ELMo, Transformers and BERT, and then details the architecture, training methods, advantages and limitations of the GPT‑1, GPT‑2, GPT‑3, ChatGPT and GPT‑4 models, concluding with a discussion of future challenges and references.

Artificial IntelligenceDeep LearningGPT
0 likes · 30 min read
Evolution of Language Models and an Overview of the GPT Series
Kuaishou Large Model
Kuaishou Large Model
Sep 27, 2023 · Artificial Intelligence

DVIS: Decoupled Framework that Sets New SOTA in Video Instance Segmentation

DVIS introduces a decoupled video instance segmentation framework that splits the task into segmentation, tracking, and refinement modules, achieving state-of-the-art performance across VIS, VPS, and VSS benchmarks while maintaining low computational overhead, and demonstrates robustness in both online and offline settings.

Deep LearningTransformercomputer vision
0 likes · 12 min read
DVIS: Decoupled Framework that Sets New SOTA in Video Instance Segmentation
DaTaobao Tech
DaTaobao Tech
Sep 27, 2023 · Artificial Intelligence

FlashAttention-2: Efficient Attention Algorithm for Transformer Acceleration and AIGC Applications

FlashAttention‑2 is an IO‑aware exact attention algorithm that cuts GPU HBM traffic through tiling and recomputation, optimizes non‑matmul FLOPs, expands sequence‑parallelism and warp‑level work distribution, delivering up to 2× speedup over FlashAttention, near‑GEMM efficiency, and enabling longer‑context Transformer training and inference for AIGC with fastunet and negligible accuracy loss.

AIGCAttention optimizationDeep Learning
0 likes · 20 min read
FlashAttention-2: Efficient Attention Algorithm for Transformer Acceleration and AIGC Applications
Kuaishou Tech
Kuaishou Tech
Sep 26, 2023 · Artificial Intelligence

Cross-Domain Product Representation (COPE): A Large-Scale Dataset and Baseline Model for Rich‑Content E‑Commerce

The paper introduces ROPE, the first large‑scale cross‑domain product recognition dataset covering detail pages, short videos and live streams, and proposes COPE, a dual‑tower multimodal model that learns unified product embeddings using contrastive and classification losses, achieving superior retrieval and few‑shot classification performance across domains.

Deep Learningcontrastive learningcross-domain
0 likes · 13 min read
Cross-Domain Product Representation (COPE): A Large-Scale Dataset and Baseline Model for Rich‑Content E‑Commerce
Kuaishou Tech
Kuaishou Tech
Sep 25, 2023 · Artificial Intelligence

LPR4M: A Large-Scale Multimodal Livestreaming Product Recognition Dataset and the RICE Cross‑View Semantic Alignment Model

This paper introduces LPR4M, a 4‑million‑pair multimodal dataset for livestreaming product recognition, and proposes the RICE model that combines instance‑level contrastive learning with patch‑level cross‑view semantic alignment, demonstrating state‑of‑the‑art performance on both LPR4M and MovingFashion benchmarks.

Deep LearningMultimodal Datasetcross-view alignment
0 likes · 19 min read
LPR4M: A Large-Scale Multimodal Livestreaming Product Recognition Dataset and the RICE Cross‑View Semantic Alignment Model
Bilibili Tech
Bilibili Tech
Sep 22, 2023 · Artificial Intelligence

AI-Based Digital Watermarking for Video: Design, Training Strategies, and Engineering Deployment

The paper presents an AI‑driven invisible video watermarking system that combines a convolutional encoder/decoder with SE blocks, a simulated‑JPEG degradation layer, multi‑term loss, block‑wise processing, anchor‑based alignment and redundancy voting, achieving high visual fidelity and robust recovery after double‑compression in large‑scale platforms like Bilibili.

AIDeep Learninganchor calibration
0 likes · 21 min read
AI-Based Digital Watermarking for Video: Design, Training Strategies, and Engineering Deployment
HomeTech
HomeTech
Sep 21, 2023 · Artificial Intelligence

Homepage Pop‑up Recommendation System for Car Purchase Intent: Background, Feature Engineering, Model and Strategy Optimization, and Results

This article details how AutoHome's homepage pop‑up leverages precise targeting, extensive feature engineering, and multi‑stage DeepFM‑based models with attention and LHUC modules to accurately identify car‑buying users, improve vehicle‑series recommendations, and achieve a 355% conversion rate increase.

AIDeep LearningMachine Learning
0 likes · 7 min read
Homepage Pop‑up Recommendation System for Car Purchase Intent: Background, Feature Engineering, Model and Strategy Optimization, and Results
Ant R&D Efficiency
Ant R&D Efficiency
Sep 19, 2023 · Artificial Intelligence

From the Turing Test to GPT‑4: A Historical Overview of Chatbots and Deep Learning

From Turing’s 1950 imitation game to GPT‑4’s multimodal vision‑language capabilities, the field has evolved from simple rule‑based programs like ELIZA and PARRY, through statistical learning and the 2017 Transformer breakthrough, to large-scale generative models that achieve fluent conversation yet still grapple with hallucination and true understanding.

Artificial IntelligenceChatbot HistoryDeep Learning
0 likes · 25 min read
From the Turing Test to GPT‑4: A Historical Overview of Chatbots and Deep Learning
Alibaba Cloud Infrastructure
Alibaba Cloud Infrastructure
Sep 13, 2023 · Artificial Intelligence

Pai‑Megatron‑Patch: Design Principles, Key Features, and End‑to‑End Usage for Large Language Model Training

This article introduces the open‑source Pai‑Megatron‑Patch tool from Alibaba Cloud, explains its non‑intrusive patch architecture, enumerates supported models and features such as weight conversion, Flash‑Attention 2.0, FP8 training with Transformer Engine, and provides detailed command‑line examples for model conversion, pre‑training, supervised fine‑tuning, inference, and RLHF reinforcement learning pipelines.

Deep LearningFP8LLM
0 likes · 19 min read
Pai‑Megatron‑Patch: Design Principles, Key Features, and End‑to‑End Usage for Large Language Model Training
NetEase Cloud Music Tech Team
NetEase Cloud Music Tech Team
Sep 6, 2023 · Artificial Intelligence

Timbre‑Guided TG‑Critic and Transformer‑Based TrOMR: AI Advances in Music Evaluation

This article reviews two recent AI research papers from NetEase Cloud Music Lab: TG‑Critic, a timbre‑guided, reference‑free singing evaluation model that classifies vocal performance using only audio, and TrOMR, a Transformer‑based end‑to‑end polyphonic optical music recognition system that improves note‑sequence prediction and dataset realism.

Audio AnalysisDeep LearningMusic Evaluation
0 likes · 6 min read
Timbre‑Guided TG‑Critic and Transformer‑Based TrOMR: AI Advances in Music Evaluation
Alibaba Cloud Developer
Alibaba Cloud Developer
Sep 4, 2023 · Artificial Intelligence

Hands‑On Building a Transformer from Scratch with PyTorch

This tutorial walks you through implementing a full Transformer model in PyTorch, starting from basic linear‑regression code, adding attention mechanisms, multi‑head attention, encoder‑decoder architecture, training loops, and inference, all reinforced with practical debugging tips.

Deep LearningNLPPyTorch
0 likes · 17 min read
Hands‑On Building a Transformer from Scratch with PyTorch
TAL Education Technology
TAL Education Technology
Aug 31, 2023 · Artificial Intelligence

Research on Content-Based Image Retrieval Techniques

This article reviews the fundamentals, feature extraction methods, evaluation metrics, and common datasets of content‑based image retrieval (CBIR), discussing traditional low‑level features, local descriptors, unsupervised and supervised learning approaches, and recent deep‑learning models for improving retrieval performance.

CBIRDeep Learningdatasets
0 likes · 13 min read
Research on Content-Based Image Retrieval Techniques
Network Intelligence Research Center (NIRC)
Network Intelligence Research Center (NIRC)
Aug 30, 2023 · Artificial Intelligence

DeepQueueNet: Scalable Network Performance Estimation with Packet‑Level Visibility

DeepQueueNet combines discrete‑event and continuous simulation with deep neural networks to deliver highly accurate, generalizable, and GPU‑scalable network performance estimates at packet‑level granularity, outperforming existing DNN‑based estimators across diverse topologies and traffic scenarios.

DESDNNDeep Learning
0 likes · 5 min read
DeepQueueNet: Scalable Network Performance Estimation with Packet‑Level Visibility
DataFunSummit
DataFunSummit
Aug 24, 2023 · Artificial Intelligence

Panoramic Indoor Layout Estimation with Vision Transformer (PanoViT)

This article introduces the PanoViT model, a vision‑transformer‑based approach for indoor layout estimation from panoramic images, covering its research background, architectural components, experimental results on public datasets, and step‑by‑step usage within ModelScope.

3D ReconstructionDeep LearningModelScope
0 likes · 8 min read
Panoramic Indoor Layout Estimation with Vision Transformer (PanoViT)
Rare Earth Juejin Tech Community
Rare Earth Juejin Tech Community
Aug 24, 2023 · Artificial Intelligence

Neural Style Transfer with PyTorch: Theory and Implementation

This article introduces neural style transfer, explains its underlying principles using VGG19 feature extraction, content and style loss definitions, and provides a complete PyTorch implementation with code for loading images, extracting features, computing Gram matrices, and optimizing the output image.

Deep LearningPyTorchStyle Transfer
0 likes · 14 min read
Neural Style Transfer with PyTorch: Theory and Implementation
Top Architect
Top Architect
Aug 22, 2023 · Artificial Intelligence

Face Recognition Search: Principles, Implementation Steps, and Applications

This article explains the background, core principles, preprocessing, feature extraction, matching algorithms, and practical application scenarios of face recognition search, and provides detailed reference implementations with Java and OpenCV code examples for building a complete system.

Deep LearningJavaOpenCV
0 likes · 15 min read
Face Recognition Search: Principles, Implementation Steps, and Applications
Ele.me Technology
Ele.me Technology
Aug 22, 2023 · Artificial Intelligence

Multi-Granularity Attention Model for Group Recommendation (MGAM)

The Multi‑Granularity Attention Model (MGAM) improves group recommendation by extracting subset, group, and superset preferences through hierarchical attention and graph neural networks, fusing them via self‑attention, and achieves state‑of‑the‑art offline results and a 1.2% online CTR lift in Alibaba’s local‑life services.

AIDeep LearningRecommendation Systems
0 likes · 18 min read
Multi-Granularity Attention Model for Group Recommendation (MGAM)
HelloTech
HelloTech
Aug 22, 2023 · Artificial Intelligence

AI Platform Architecture and Automation in Machine Learning

An end‑to‑end AI platform integrates feature processing, model training, deployment, and decision orchestration across offline and online layers, leveraging automated pipelines such as AutoML (feature engineering, hyper‑parameter optimization, neural architecture search) built on Ray Tune and NNI, which have already boosted CTR in real‑world advertising and aim to make every user an algorithm engineer.

AI platformAutoMLDeep Learning
0 likes · 8 min read
AI Platform Architecture and Automation in Machine Learning
DaTaobao Tech
DaTaobao Tech
Aug 21, 2023 · Artificial Intelligence

Action Sensitivity Learning for Temporal Action Localization

The paper presents Action Sensitivity Learning (ASL), a framework that models frame‑wise importance at both class‑level (via learnable Gaussian distributions) and instance‑level (using quality scores), integrates these weights into classification and regression losses, adds a contrastive InfoNCE term, and achieves state‑of‑the‑art temporal action localization performance across six benchmark datasets.

Action Sensitivity LearningDeep LearningTemporal Action Localization
0 likes · 8 min read
Action Sensitivity Learning for Temporal Action Localization
Ele.me Technology
Ele.me Technology
Aug 16, 2023 · Artificial Intelligence

Spatiotemporal-Enhanced Network for Click-Through Rate Prediction in Location‑Based Services

The paper introduces StEN, a spatiotemporal-enhanced network for CTR prediction in location-based services, combining static spatiotemporal feature activation, dynamic preference activation, and target attention, achieving state-of-the-art offline results and a 1.6% CTR lift in online tests.

Deep LearningRecommendation SystemsSpatiotemporal Modeling
0 likes · 19 min read
Spatiotemporal-Enhanced Network for Click-Through Rate Prediction in Location‑Based Services
Rare Earth Juejin Tech Community
Rare Earth Juejin Tech Community
Aug 16, 2023 · Artificial Intelligence

Deep Dive into OCR – Chapter 2: Development and Classification of OCR Technology

This article provides a comprehensive overview of OCR technology, detailing the evolution from traditional hand‑crafted methods to modern deep‑learning approaches, describing image preprocessing, text detection and recognition pipelines, summarizing classic machine‑learning algorithms, and presenting a practical OpenCV implementation with Python code.

Deep LearningOCROpenCV
0 likes · 23 min read
Deep Dive into OCR – Chapter 2: Development and Classification of OCR Technology
Rare Earth Juejin Tech Community
Rare Earth Juejin Tech Community
Aug 12, 2023 · Artificial Intelligence

An Introduction to OCR: Concepts, History, Applications, Datasets, and Technical Workflow

This article provides a comprehensive overview of Optical Character Recognition (OCR), covering its definition, historical development, classification, real‑world applications, technical pipeline, common challenges, mitigation strategies, popular datasets, model performance comparisons, and leading open‑source platforms.

Deep LearningOCROptical Character Recognition
0 likes · 16 min read
An Introduction to OCR: Concepts, History, Applications, Datasets, and Technical Workflow
Kuaishou Tech
Kuaishou Tech
Aug 11, 2023 · Artificial Intelligence

PEPNet: Parameter and Embedding Personalized Network for Multi‑Task Multi‑Domain Recommendation

The paper introduces PEPNet, a plug‑and‑play network that tackles the domain‑seesaw and task‑seesaw problems in multi‑scenario recommendation by using a gated personalization module (GateNU) together with embedding‑level (EPNet) and parameter‑level (PPNet) personalization, and demonstrates its superiority through extensive offline and online experiments on Kuaishou data.

Deep Learningembeddinggate network
0 likes · 11 min read
PEPNet: Parameter and Embedding Personalized Network for Multi‑Task Multi‑Domain Recommendation
Rare Earth Juejin Tech Community
Rare Earth Juejin Tech Community
Jul 31, 2023 · Artificial Intelligence

Overview of Deep Neural Network Architectures

This article provides a comprehensive overview of deep neural network families, introducing twelve major architectures—including Feedforward, CNN, RNN, LSTM, DBN, GAN, Autoencoder, Residual, Capsule, Transformer, Attention, and Deep Reinforcement Learning—explaining their principles, structures, training methods, and offering Python/TensorFlow/PyTorch code examples.

CNNDeep LearningGaN
0 likes · 29 min read
Overview of Deep Neural Network Architectures
Smart Era Software Development
Smart Era Software Development
Jul 26, 2023 · Artificial Intelligence

Three Core Skills Every Aspiring AI Architect Needs

The article defines the AI architect role and outlines three essential capabilities—mastery of AI technologies and development workflow, deep business understanding with strong abstraction ability, and the design and implementation of efficient, scalable AI solutions—explaining why each is critical for successful AI product delivery.

AI architectureDeep LearningMachine Learning
0 likes · 10 min read
Three Core Skills Every Aspiring AI Architect Needs