Tagged articles

On-Device Inference

10 articles · Page 1 of 1
Geek Labs
Geek Labs
Sep 28, 2026 · Artificial Intelligence

OpenSquilla: Local Routing Slashes AI Agent Costs 9x Without Quality Loss

OpenSquilla, a 7K-star open-source AI agent, uses on-device routing to classify each conversation turn by complexity and dispatch it to the cheapest suitable model, achieving 9x cost reduction on 25 benchmark tasks while maintaining near-identical scores, plus adaptive reasoning, dynamic prompts, and pluggable providers.

AI AgentBenchmarkCost Optimization
0 likes · 9 min read
OpenSquilla: Local Routing Slashes AI Agent Costs 9x Without Quality Loss
macrozheng
macrozheng
Sep 23, 2026 · Mobile Development

Run LLMs on 6GB Android Phones with OlliteRT: OpenAI-Compatible LAN API

OlliteRT is an open-source Android app that runs large language models locally via Google's LiteRT runtime, exposing an OpenAI-compatible API over LAN for lightweight always-on tasks like RSS summarization and smart home automation, with monitoring, logging, and security features.

AndroidEdge AIGemma
0 likes · 11 min read
Run LLMs on 6GB Android Phones with OlliteRT: OpenAI-Compatible LAN API
51CTO HarmonyOS Developer Community
51CTO HarmonyOS Developer Community
Jul 13, 2026 · Mobile Development

HarmonyOS 7 Local LLM Integration: Capabilities, Benchmarks & Engineering Guide

This article details integrating local LLMs into HarmonyOS apps, covering use cases like narrative generation and offline privacy, real-world benchmarks on Mate 60 Pro with Qwen2.5-0.5B, architecture using ArkTS and llama.cpp, performance optimizations via O3/LTO/KleidiAI, and key pitfalls like token limits and UI threading.

GGUFHarmonyOSLocal LLM
0 likes · 15 min read
HarmonyOS 7 Local LLM Integration: Capabilities, Benchmarks & Engineering Guide
Sohu Tech Products
Sohu Tech Products
Nov 5, 2025 · Artificial Intelligence

How nndeploy Simplifies the Last Mile of On-Device AI Deployment

nndeploy is an open‑source, high‑performance on‑device AI deployment framework that abstracts the repetitive “last‑mile” workflow into a visual drag‑and‑drop DAG, offering multi‑platform inference, optimization, and ready‑to‑use model configs, enabling developers to go from prototype to production in minutes.

AI DeploymentEdge AIOn-Device Inference
0 likes · 15 min read
How nndeploy Simplifies the Last Mile of On-Device AI Deployment
DaTaobao Tech
DaTaobao Tech
Oct 14, 2024 · Artificial Intelligence

MNN Stable Diffusion: On‑Device Deployment and Performance Optimizations

The article presents Alibaba’s open‑source MNN inference engine, demonstrating how quantization, operator fusion (including fused multi‑head attention, GroupNorm/SplitGeLU, Winograd convolutions), optimized GEMM and memory‑paging enable on‑device Stable Diffusion with 1‑second‑per‑step performance on Snapdragon 8 Gen3 and Apple M3 GPUs, and outlines future speed‑up directions.

AIMNNOn-Device Inference
0 likes · 11 min read
MNN Stable Diffusion: On‑Device Deployment and Performance Optimizations
Sohu Tech Products
Sohu Tech Products
Mar 6, 2024 · Mobile Development

On‑Device Deployment of Large Language Models Using Sohu’s Hybrid AI Engine and GPT‑2

The article outlines how Sohu’s Hybrid AI Engine enables on‑device deployment of a distilled GPT‑2 model by converting it to TensorFlow Lite, detailing the setup, customization with Keras, inference workflow, and core SDK calls, and argues that this approach offers fast, private, and cost‑effective AI for mobile devices despite typical LLM constraints.

GPT-2Hybrid AIKeras
0 likes · 9 min read
On‑Device Deployment of Large Language Models Using Sohu’s Hybrid AI Engine and GPT‑2
DataFunTalk
DataFunTalk
Feb 10, 2022 · Artificial Intelligence

Evolution of Re‑ranking Techniques in Kuaishou Short‑Video Recommendation System

This article details the technical evolution of Kuaishou's short‑video recommendation pipeline, focusing on sequence re‑ranking, multi‑content mixing, and on‑device re‑ranking, and explains how transformer‑based models, generator‑evaluator frameworks, and reinforcement‑learning strategies are employed to maximize overall sequence value, user engagement, and revenue.

KuaishouOn-Device InferenceReinforcement Learning
0 likes · 15 min read
Evolution of Re‑ranking Techniques in Kuaishou Short‑Video Recommendation System
Sohu Tech Products
Sohu Tech Products
Jan 20, 2021 · Mobile Development

Hybrid AI Engine: Integrating On‑Device Image Recognition with TensorFlow Lite and HiAI

This article introduces three traditional approaches for deploying machine‑learning models on mobile devices, analyzes their drawbacks, and presents a hybrid AI engine that combines TensorFlow Lite and system‑level HiAI to provide a unified, lightweight, and developer‑friendly on‑device image‑recognition solution, including code examples.

Android developmentHybrid AI EngineOn-Device Inference
0 likes · 12 min read
Hybrid AI Engine: Integrating On‑Device Image Recognition with TensorFlow Lite and HiAI