Tagged articles

batch inference

2 articles · Page 1 of 1
Fei's Miscellaneous Talks
Fei's Miscellaneous Talks
Sep 6, 2026 · Artificial Intelligence

How Prediction Servers Score 1000 Candidates in Milliseconds: Fine-Ranking Architecture Deep Dive

This article details the architecture and optimization of a Prediction Server for fine-ranking in recommender systems, covering model management, batch inference, multi-objective fusion, probability calibration, deployment pipelines, and high-availability patterns to score thousands of candidates within milliseconds.

ONNX Runtimebatch inferencefine ranking
0 likes · 46 min read
How Prediction Servers Score 1000 Candidates in Milliseconds: Fine-Ranking Architecture Deep Dive
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Jul 31, 2026 · Big Data

Dual‑Dimension Cost Cutting for EMR Serverless Spark AI Functions

The article explains how EMR Serverless Spark AI Functions incur costs from model inference and Spark compute, and presents a two‑pronged cost‑saving strategy—AI query optimization to cut unnecessary calls and asynchronous Batch File inference to lower unit prices and release executor resources—complete with examples, benchmarks, and configuration guidance.

AI FunctionCost OptimizationEMR Serverless
0 likes · 20 min read
Dual‑Dimension Cost Cutting for EMR Serverless Spark AI Functions