Tagged articles

Inference Engine

20 articles · Page 1 of 1
Architecture Digest
Architecture Digest
Sep 21, 2026 · Artificial Intelligence

3 AI Open-Source Projects: DeepSeek Agent Runtime, Verified Diagrams, 744B MoE on 25GB RAM

This article reviews three cutting-edge AI open-source projects: DeepSeek's plugin-based Agent runtime framework (deepseek-harness), archify for generating verifiable architecture diagrams directly in coding agents, and colibri, a pure C inference engine enabling 744B MoE models to run on 25GB RAM via disk streaming.

Agent FrameworkArchitecture DiagramsDeepSeek
0 likes · 8 min read
3 AI Open-Source Projects: DeepSeek Agent Runtime, Verified Diagrams, 744B MoE on 25GB RAM
AI Architecture Path
AI Architecture Path
Sep 17, 2026 · Artificial Intelligence

Run 744B MoE Model on 25GB RAM: Colibri's Tiered Storage Breakthrough

Colibri, a pure C inference engine with zero dependencies, enables running the 744B parameter GLM-5.2 MoE model on consumer hardware with just 25GB RAM and NVMe SSD by leveraging MoE sparsity and a three-tier storage scheduling system across VRAM, RAM, and disk, achieving 0.05–6.8 tok/s depending on hardware.

ColibriGLM-5.2Inference Engine
0 likes · 14 min read
Run 744B MoE Model on 25GB RAM: Colibri's Tiered Storage Breakthrough
Architecture Digest
Architecture Digest
Sep 13, 2026 · Artificial Intelligence

Running 2.78T-Parameter Kimi K3 on 8GB RAM: Pure C Inference Engine Explained

A pure C inference engine runs the 2.78 trillion parameter Kimi K3 model on just 8GB RAM by streaming weights from disk, leveraging MoE sparsity (only 3.7% active per token) and computing directly on compressed formats, achieving correct output at 32.69 seconds per token on consumer hardware.

C languageDisk StreamingInference Engine
0 likes · 11 min read
Running 2.78T-Parameter Kimi K3 on 8GB RAM: Pure C Inference Engine Explained
AI Large-Model Wave and Transformation Guide
AI Large-Model Wave and Transformation Guide
Sep 11, 2026 · Industry Insights

Why Material Issuance ≠ Cost Transfer: OntoL's Ontology Modeling Fixes Business-Finance Disconnect

OntoL uses dual ontology modeling — separating business material flow from financial cost recognition — to decouple warehouse issuance from profit-and-loss impact, enabling automatic WIP tracking, real-time discrepancy detection, and full audit traceability, cutting month-end closing from seven days to half a day in manufacturing case studies.

Inference EngineTBox/ABoxWIP cost accounting
0 likes · 13 min read
Why Material Issuance ≠ Cost Transfer: OntoL's Ontology Modeling Fixes Business-Finance Disconnect
AI Engineer Programming
AI Engineer Programming
Aug 21, 2026 · Artificial Intelligence

Essential Concepts and Terminology for Deploying Large Language Models Locally

This article walks through the core concepts needed before deploying a large language model on‑premises, covering weight precision, quantization methods, model packaging formats, inference engines, GPU memory considerations, KV‑cache sizing, sampling strategies, optional extensions such as LoRA and RAG, and a step‑by‑step decision workflow to match hardware, model, and deployment goals.

Inference EngineKV CacheLLM
0 likes · 21 min read
Essential Concepts and Terminology for Deploying Large Language Models Locally
ThinkingAgent
ThinkingAgent
Jul 5, 2026 · Artificial Intelligence

Building L1 AI Infra: Model Gateways, Smart Routing, and High‑Performance Inference Engines

The article presents a comprehensive, production‑ready guide for the L1 layer of AI infrastructure, detailing how model gateways unify calls, intelligent routing selects the optimal model, inference engines maximize GPU throughput, and quantization and KV‑Cache techniques dramatically cut costs while maintaining performance.

AI infrastructureInference EngineKV Cache
0 likes · 25 min read
Building L1 AI Infra: Model Gateways, Smart Routing, and High‑Performance Inference Engines
AI Large-Model Wave and Transformation Guide
AI Large-Model Wave and Transformation Guide
Jun 8, 2026 · Artificial Intelligence

Designing a High‑Reliability Cognitive Reasoning System with Ontology‑Based Architecture

The article presents a detailed architecture for a high‑reliability cognitive reasoning system that combines logical inference, semantic constraints, and a seven‑layer defense to achieve efficient deduction and strict error prevention across critical domains such as medical diagnosis and financial risk control.

Inference Enginecognitive reasoningexplainable AI
0 likes · 6 min read
Designing a High‑Reliability Cognitive Reasoning System with Ontology‑Based Architecture
DeepHub IMBA
DeepHub IMBA
Apr 4, 2026 · Artificial Intelligence

Building Mini-vLLM from Scratch: KV‑Cache, Dynamic Batching, and Distributed Inference

This article walks through constructing Mini-vLLM, a from‑scratch LLM inference engine that tackles the O(N²) attention cost with KV‑cache, boosts throughput via dynamic batching, adds observability with Prometheus/Grafana, supports gRPC, and scales across multiple workers, with benchmark numbers demonstrating its CPU‑only performance.

DockerInference EngineKV Cache
0 likes · 12 min read
Building Mini-vLLM from Scratch: KV‑Cache, Dynamic Batching, and Distributed Inference
DataFunSummit
DataFunSummit
Dec 24, 2024 · Artificial Intelligence

Considerations and Practices for Domesticating Large‑Model Inference Engines

This article examines the importance of domestic large‑model inference engines, compares Chinese and international chips, evaluates four architectural approaches, discusses practical challenges such as performance loss and model support, and outlines future expectations for high‑performance, heterogeneous‑chip inference solutions.

Inference EnginePerformance Optimizationdomestic chip
0 likes · 9 min read
Considerations and Practices for Domesticating Large‑Model Inference Engines
DataFunSummit
DataFunSummit
Sep 11, 2023 · Artificial Intelligence

Challenges and Insights for Deploying Large Models on Edge with MNN

The talk presents an overview of the MNN inference engine, outlines the end‑to‑end workflow for deploying large language models on mobile devices, discusses technical challenges and practical solutions, and concludes with future directions for edge AI deployment.

AIEdge DeploymentInference Engine
0 likes · 2 min read
Challenges and Insights for Deploying Large Models on Edge with MNN
OPPO Kernel Craftsman
OPPO Kernel Craftsman
Oct 28, 2022 · Artificial Intelligence

ShaderNN: A GPU Shader‑Based Lightweight Inference Engine for Mobile AI Applications

ShaderNN is an open‑source, sub‑2 MB GPU‑shader inference engine that runs TensorFlow, PyTorch and ONNX models directly on mobile graphics textures via OpenGL fragment and compute shaders, delivering real‑time, low‑power AI for image‑heavy tasks while eliminating third‑party dependencies and achieving up to 90 % speed gains.

GPUInference EnginePerformance
0 likes · 11 min read
ShaderNN: A GPU Shader‑Based Lightweight Inference Engine for Mobile AI Applications
ByteDance Terminal Technology
ByteDance Terminal Technology
Jul 29, 2022 · Artificial Intelligence

Pitaya: ByteDance’s End‑Side AI Engineering Platform Overview

Pitaya, built by ByteDance’s Client AI and MLX teams, is a comprehensive end‑side AI engineering platform that provides a full workflow from model development and data preparation to deployment, monitoring, and federated learning, supporting large‑scale commercial scenarios across multiple apps.

AI platformEdge AIInference Engine
0 likes · 14 min read
Pitaya: ByteDance’s End‑Side AI Engineering Platform Overview
DataFunTalk
DataFunTalk
Apr 14, 2022 · Artificial Intelligence

PaddlePaddle Deep Learning Platform: Architecture, Core Technologies, and Real‑World Applications

The article presents a comprehensive overview of Baidu's open‑source deep learning platform PaddlePaddle, detailing its full‑stack architecture, core technologies such as unified dynamic‑static graph, large‑scale distributed training, multi‑platform inference, an extensive model zoo, hardware adaptation, and showcases a real‑world deployment case in power‑grid monitoring.

AI FrameworkDistributed TrainingInference Engine
0 likes · 15 min read
PaddlePaddle Deep Learning Platform: Architecture, Core Technologies, and Real‑World Applications
DaTaobao Tech
DaTaobao Tech
Mar 11, 2022 · Artificial Intelligence

How Alibaba’s MNN Engine Achieves 350% CPU Speedup and Sparse Acceleration

Alibaba’s MNN, a lightweight high‑performance deep‑learning inference engine, earned top honors in China’s 2022 “Science & Innovation China” awards, and delivers impressive gains such as 350% speedup on X86 CPUs, 2.1‑2.3× acceleration on ARM with sparse models, plus integrated OpenCV/Numpy functionality for edge AI deployment.

AI DeploymentAlibabaInference Engine
0 likes · 4 min read
How Alibaba’s MNN Engine Achieves 350% CPU Speedup and Sparse Acceleration
Alibaba Terminal Technology
Alibaba Terminal Technology
Feb 3, 2021 · Frontend Development

How Front-End AI Inference Engines Achieve Real-Time Smart Recognition

This article explains on‑device machine learning concepts, compares front‑end inference engines such as TensorFlow.js, ONNX.js and WebDNN across CPU, WASM and WebGL, and presents practical optimization techniques like vectorization, memory layout, graph fusion and mixed‑precision to boost performance for real‑time applications.

Inference Enginefrontendmachine learning
0 likes · 11 min read
How Front-End AI Inference Engines Achieve Real-Time Smart Recognition
Alibaba Cloud Developer
Alibaba Cloud Developer
Jul 2, 2019 · Artificial Intelligence

How MNN Powers Mobile AI: Inside Alibaba’s Open‑Source Inference Engine

Alibaba’s MNN (Mobile Neural Network) engine, now open‑sourced on GitHub, showcases how a lightweight, end‑side deep‑learning inference framework tackles fragmentation, optimizes model conversion, scheduling, and execution across diverse devices, delivering significant performance gains for mobile and IoT AI applications.

Inference EngineMNNgraph optimization
0 likes · 15 min read
How MNN Powers Mobile AI: Inside Alibaba’s Open‑Source Inference Engine